Uber Built Custom Controller to Cut Compute Overhead

The firm reduced provisioning requirements by decoupling workload orchestration across its global infrastructure.

Updated on Sept. 28, 2026 in Data Centers

Isometric editorial illustration featuring a central vertical monolith surrounded by stacked geometric compute blocks in shades of teal, slate, and cream.
Uber has deployed its custom ServiceScale controller, a new internal architecture that successfully reduced the company’s compute provisioning overhead by 35%. AI Illustration. Upload story photo >

Live Poll

Do you trust that complex automated systems can reliably manage infrastructure failovers without human intervention?

Uber has deployed a custom ServiceScale controller to manage workload orchestration across its compute clusters. This internal architecture has successfully reduced the company’s steady-state provisioning multiplier from 2x to 1.3x.

Why it matters

By decoupling scaling intent from execution, the firm eliminated over one million active CPU cores without impacting failover reliability. The change allows Uber to repurpose idle capacity from low-tier workloads for regional failover operations.

The system manages 1.5 million daily pod launches across 4,000 services and 100 clusters. It successfully dropped provisioning overhead to 1.3x while eliminating one million CPU cores.

The players

Uber

A global transportation and logistics company that manages 3 million CPU cores and 100 compute clusters to maintain its distributed services stack.

The details

The ServiceScale controller functions by using a custom resource definition—a K8s object extension that enables users to define unique API types—to reconcile scaling desires into native Kubernetes objects. The team introduced read-your-own-write consistency guardrails using generation annotations to prevent metadata drift. Additionally, an automated healer monitors the fleet to resolve ReplicaSet metadata discrepancies between multiple cloud providers.

Timeline

  1. January 2026: Uber published a paper on its failover architecture.

  2. April 2026: Kubernetes v1.36 was released featuring staleness mitigation.

  3. September 28, 2026: Uber published the account of its ServiceScale controller.

The Tech Race

Uber's implementation sits alongside efforts to optimize multi-cloud orchestration for massive-scale distributed services. It marks a significant departure from standard controller behaviors by prioritizing regional failover efficiency over static capacity reserves.

This architecture optimizes resource allocation for backend compute systems rather than consumer-facing application interfaces. Engineering teams managing massive distributed services may look to these architectural patterns to reduce their own idle cloud capacity costs.

The takeaway

Uber’s approach demonstrates how custom metadata handling can solve the long-standing industry challenge of over-provisioning for failover events. Observers should track Kubernetes API developments that may eventually standardize these custom resource reconciliation methods.

Further reading

For broader trends in large-scale infrastructure management, see our Data Centers section.

Source note: This article includes information reported by InfoQ.

Live Poll

Do you trust that complex automated systems can reliably manage infrastructure failovers without human intervention?