Clockwork.io Raised $31 Million for AI Fault Tolerance

The startup aims to minimize GPU downtime for large-scale training clusters through new checkpointing solutions.

Updated on Oct. 5, 2026 in Artificial Intelligence

Clockwork.io Raised $31 Million for AI Fault Tolerance

Live Poll

Do you believe investments in AI infrastructure reliability are necessary for modern business performance?

Clockwork.io has raised $31 million in new funding to expand its fault-tolerance software suite for AI infrastructure. The company provides tools to manage GPU failures and network interruptions in large-scale distributed training environments.

Why it matters

Distributed AI workloads are highly susceptible to hardware and network failures, which can derail training runs for days. By automating workload migration and traffic rerouting, Clockwork.io aims to improve compute efficiency and reliability for massive GPU deployments.

Clockwork.io reported that its TorchPass solution reduced training goodput loss from 14% to 3% in neocloud testing. The platform supports asynchronous application checkpoints to mitigate downtime during large-scale operations.

The players

Clockwork.io

A Palo Alto-based firm developing software for fault tolerance in AI training infrastructure.

LinkedIn

A professional social networking platform currently deploying LinkPass across its infrastructure fleet.

Together AI

An AI compute provider that offers TorchPass as a service on its dedicated GPU clusters.

Premji Invest

A global investment firm that co-led the most recent $31 million funding round.

The details

The platform addresses distributed system failures using two primary mechanisms: LinkPass and TorchPass. LinkPass functions by dynamically rerouting network traffic around failed links to maintain connectivity during infrastructure issues. TorchPass facilitates workload migration by moving active training jobs from a failing GPU to a healthy one, helping to avoid total job restart requirements.

Timeline

  1. 2021: NEA provided initial venture backing for the company.

  2. October 5, 2026: The company announced the successful close of its $31 million funding round.

The Tech Race

Clockwork.io is competing to solve the stability bottleneck inherent in training models across thousands of GPUs, a challenge highlighted by the 16,384-unit cluster scale used in Meta Llama 3 training. The company follows a trend of increasing infrastructure software investment as commercial AI providers prioritize uptime to remain competitive.

Enterprise AI users and cloud providers can currently implement these tools to reduce training goodput loss significantly. While the platform is available through partners like Together AI, widespread adoption depends on its integration into broader data center management workflows.

The takeaway

Reliability software is becoming a critical layer for developers managing massive compute clusters that would otherwise suffer from frequent downtime. Watch for further adoption figures among major cloud providers as a key indicator of whether these fault-tolerance tools become a standard in the AI training stack.

Further reading

For more on the challenges of large-scale distributed computing, visit the Artificial Intelligence section.

More information

Learn more about the platform architecture on the Clockwork.io official company website.

Live Poll

Do you believe investments in AI infrastructure reliability are necessary for modern business performance?