For AI agents, 'the run completed' isn't success. Lemma raises $2.3M to prove it
The Y Combinator Fall 2025 alum, founded by Jerry Zhang and Cole Gawin, drew a broad syndicate including Matrix, Liquid 2 Ventures, Cervin Ventures and angels from OpenAI, xAI, Meta and DoorDash.
Lemma, a San Francisco startup building production monitoring for AI agents, has raised $2.3M in pre-seed financing, the company announced in August 2026. The round backs a problem that becomes more expensive as AI agents move from demonstrations into real workflows: traditional monitoring can show that software executed, but an agent can complete every technical step and still misunderstand the assignment, make the wrong decision, or quietly produce an unusable outcome.
Founded in 2025 by Jerry Zhang and Cole Gawin, Lemma participated in Y Combinator's Fall 2025 batch. The company says its platform analyzes production traces against intended behavior, groups recurring failures, alerts teams, carries incident context into coding workflows, and monitors for regressions after fixes. The larger question, as the company frames it, is no longer whether an agent ran; it is whether the agent did the right work.
A funding summary reposted by Lemma names Matrix, Y Combinator, Liquid 2 Ventures, Vermilion Cliffs Ventures, Irregular Expressions, Cervin Ventures, Comma Capital, Position Ventures, Eight Capital, and angels affiliated with OpenAI, xAI, Meta, and DoorDash as participants. The reviewed sources do not identify a lead investor, valuation, security structure, or formal allocation for the proceeds. The clean version is still meaningful: Lemma has $2.3M in new capital and a broad early-stage syndicate behind its attempt to make AI-agent behavior measurable in production.
The round arrives while the company is visibly expanding. Lemma's official site lists teams including Modus, Boardy, Stan, Folk, and Hyperspell under "Trusted by teams shipping agents," while its Y Combinator profile shows open roles across engineering, product design, and developer relations. Those are company-presented adoption and hiring signals, not independently audited revenue or retention metrics.
Conventional observability was built for systems where failure usually leaves technical evidence — a failed test, an exception, a service that stops responding, a value outside a threshold. Those signals remain useful, but they do not fully describe an AI agent whose software stack ran cleanly while its reasoning produced the wrong result. That is the failure class Lemma targets: comparing an agent's behavior with the instructions and context that define success, then surfacing issues a team did not know to encode as a rule in advance. Its website describes a workflow that moves from trace analysis to issue grouping, Slack alerts, coding-agent context, and online evaluation after a fix.
Lemma reports monitoring more than 1M agent traces per day. That figure is company-reported and has not been independently audited, but it offers a picture of the product's intended operating level. The platform is not framed as a lab notebook for occasional experiments; it is built for recurring production traffic where rare failures become routine at scale.
Press
About the Company
Catch AI agents failing in production