Clockwork.io Raised $31 Million for AI Fault Tolerance
The startup aims to minimize GPU downtime for large-scale training clusters through new checkpointing solutions.
Updated on Oct. 5, 2026 in Artificial Intelligence

Live Poll
Do you believe investments in AI infrastructure reliability are necessary for modern business performance?
Clockwork.io has raised $31 million in new funding to expand its fault-tolerance software suite for AI infrastructure. The company provides tools to manage GPU failures and network interruptions in large-scale distributed training environments.
Why it matters
Distributed AI workloads are highly susceptible to hardware and network failures, which can derail training runs for days. By automating workload migration and traffic rerouting, Clockwork.io aims to improve compute efficiency and reliability for massive GPU deployments.
Clockwork.io reported that its TorchPass solution reduced training goodput loss from 14% to 3% in neocloud testing. The platform supports asynchronous application checkpoints to mitigate downtime during large-scale operations.
The players
Clockwork.io
A Palo Alto-based firm developing software for fault tolerance in AI training infrastructure.
A professional social networking platform currently deploying LinkPass across its infrastructure fleet.
Together AI
An AI compute provider that offers TorchPass as a service on its dedicated GPU clusters.
Premji Invest
A global investment firm that co-led the most recent $31 million funding round.
The details
The platform addresses distributed system failures using two primary mechanisms: LinkPass and TorchPass. LinkPass functions by dynamically rerouting network traffic around failed links to maintain connectivity during infrastructure issues. TorchPass facilitates workload migration by moving active training jobs from a failing GPU to a healthy one, helping to avoid total job restart requirements.
Timeline
2021: NEA provided initial venture backing for the company.
October 5, 2026: The company announced the successful close of its $31 million funding round.
The Tech Race
Clockwork.io is competing to solve the stability bottleneck inherent in training models across thousands of GPUs, a challenge highlighted by the 16,384-unit cluster scale used in Meta Llama 3 training. The company follows a trend of increasing infrastructure software investment as commercial AI providers prioritize uptime to remain competitive.
Enterprise AI users and cloud providers can currently implement these tools to reduce training goodput loss significantly. While the platform is available through partners like Together AI, widespread adoption depends on its integration into broader data center management workflows.
The takeaway
Reliability software is becoming a critical layer for developers managing massive compute clusters that would otherwise suffer from frequent downtime. Watch for further adoption figures among major cloud providers as a key indicator of whether these fault-tolerance tools become a standard in the AI training stack.
Further reading
For more on the challenges of large-scale distributed computing, visit the Artificial Intelligence section.
More information
Learn more about the platform architecture on the Clockwork.io official company website.
Live Poll
Do you believe investments in AI infrastructure reliability are necessary for modern business performance?








