Clockwork.io's $31M bet: keeping GPU clusters productive when hardware fails
Premji Invest, Wing Venture Capital and Seligman Ventures co-led the financing, which brings Clockwork's total funding to $73M and arrives with production deployments at LinkedIn and Together AI.
Clockwork.io has raised $31M to keep distributed AI workloads running when the hardware underneath them does not. The Palo Alto company builds software for observability, fault tolerance and performance optimization across GPU clusters, and it disclosed the financing on October 5, 2026, without a round label, valuation, investor allocations or ownership terms. The transaction is best described as new venture funding rather than forced into an unsupported series.
Premji Invest, Wing Venture Capital and Seligman Ventures co-led. Existing backers NEA and e& Capital participated. Clockwork says the $31M brings its total funding to $73M; it previously announced a $21M Series A led by NEA in March 2022, and the $73M total implies additional financing before this round without a complete, primary-source breakdown of every prior transaction.
The company said it will use the capital to accelerate its fault-tolerance suite across training, inference and reinforcement learning, expand enterprise adoption and scale delivery through cloud partners.
Why that layer matters is an economic question. Large AI workloads coordinate thousands of GPUs as one distributed system, and a failed accelerator, cable, NIC, link or server can stall progress beyond the damaged component because the remaining devices still have to reach the same synchronization points. The bill continues while healthy hardware waits or repeats computation lost since the last checkpoint. Meta's Llama 3 research offers a scale reference: during a 54-day snapshot of Llama 3 405B pre-training on 16,384 H100 GPUs, Meta recorded 419 unexpected interruptions, roughly one every three hours, while still reporting more than 90% effective training time. Clockwork is selling against that gap, framing fault tolerance as part of compute economics rather than a maintenance tool deployed after an outage. The metric it points to is goodput — the share of GPU time that advances the workload instead of waiting, recovering or repeating completed work.
The product grew from software-based clock synchronization research at Stanford, which gives the system fine-grained visibility into network and workload behavior. LinkPass reroutes traffic around a failed network path so a link problem does not interrupt a running job. TorchPass moves work from a failing GPU to a healthy one instead of rolling back to an earlier checkpoint. Clockwork also introduced multi-node platform snapshots that capture the state of a distributed job across every node, and asynchronous application checkpoints for reinforcement learning that are designed to capture updated model weights while a job continues and move them to inference replicas generating new rollouts.
Adoption gives the round its commercial shape. LinkedIn says it has deployed LinkPass across its AI infrastructure fleet; Raghu Hiremagalur, LinkedIn's SVP and CTO of Infrastructure, said the software reroutes traffic during link, optic, cable and NIC failures and prevents tens of thousands of GPU-hours of downtime each month — an attributed customer figure published in Clockwork's release, not an independently audited metric. Together AI is bringing TorchPass to market as a service on its GPU clusters and plans to demonstrate a multi-node training job continuing through injected network and GPU failures. WhiteFiber is expanding Clockwork software across its GPU-as-a-service footprint, covering an enterprise operator, an AI cloud platform and a GPU infrastructure provider.
About the Company
Turns precise timing into GPU observability and fault tolerance, helping teams keep distributed workloads moving through infrastructure failures; used by LinkedIn, Together AI and WhiteFiber, bringing total funding to $73M.