Uber Built Custom Controller to Cut Compute Overhead
The firm reduced provisioning requirements by decoupling workload orchestration across its global infrastructure.
Updated on Sept. 28, 2026 in Data Centers

Live Poll
Do you trust that complex automated systems can reliably manage infrastructure failovers without human intervention?
Uber has deployed a custom ServiceScale controller to manage workload orchestration across its compute clusters. This internal architecture has successfully reduced the company’s steady-state provisioning multiplier from 2x to 1.3x.
Why it matters
By decoupling scaling intent from execution, the firm eliminated over one million active CPU cores without impacting failover reliability. The change allows Uber to repurpose idle capacity from low-tier workloads for regional failover operations.
The system manages 1.5 million daily pod launches across 4,000 services and 100 clusters. It successfully dropped provisioning overhead to 1.3x while eliminating one million CPU cores.
The players
Uber
A global transportation and logistics company that manages 3 million CPU cores and 100 compute clusters to maintain its distributed services stack.
The details
The ServiceScale controller functions by using a custom resource definition—a K8s object extension that enables users to define unique API types—to reconcile scaling desires into native Kubernetes objects. The team introduced read-your-own-write consistency guardrails using generation annotations to prevent metadata drift. Additionally, an automated healer monitors the fleet to resolve ReplicaSet metadata discrepancies between multiple cloud providers.
Timeline
January 2026: Uber published a paper on its failover architecture.
April 2026: Kubernetes v1.36 was released featuring staleness mitigation.
September 28, 2026: Uber published the account of its ServiceScale controller.
The Tech Race
Uber's implementation sits alongside efforts to optimize multi-cloud orchestration for massive-scale distributed services. It marks a significant departure from standard controller behaviors by prioritizing regional failover efficiency over static capacity reserves.
This architecture optimizes resource allocation for backend compute systems rather than consumer-facing application interfaces. Engineering teams managing massive distributed services may look to these architectural patterns to reduce their own idle cloud capacity costs.
The takeaway
Uber’s approach demonstrates how custom metadata handling can solve the long-standing industry challenge of over-provisioning for failover events. Observers should track Kubernetes API developments that may eventually standardize these custom resource reconciliation methods.
Further reading
For broader trends in large-scale infrastructure management, see our Data Centers section.
Source note: This article includes information reported by InfoQ.
Live Poll
Do you trust that complex automated systems can reliably manage infrastructure failovers without human intervention?









