Ocular AI's $2M pre-seed bets voice models need better evidence, not just more data
Drive Capital led the round, with Y Combinator, Alumni Ventures, 1745 Ventures, Orange Collective and MyAsia VC participating; the cash funds more high-fidelity datasets, evaluations and engineering and research hires.
Ocular AI, a San Francisco startup building training data and evaluation benchmarks for voice and multimodal models, announced a $2M pre-seed round on October 5, 2026, led by Drive Capital. Y Combinator, Alumni Ventures, 1745 Ventures, Orange Collective, MyAsia VC and a group of unnamed angel investors also participated. The company did not disclose its valuation or the date the financing closed.
The pitch rests on a gap Ocular says widens as models improve: the internet holds enormous amounts of speech and video, but relatively little data that preserves the timing, overlap, expertise, consent and context real products need. A transcript can preserve every word and still discard much of a conversation. People overlap, interrupt, hesitate, restart sentences, change emphasis, and acknowledge one another with short sounds, facial expressions or gestures.
Ocular captures full-duplex conversations with separate high-fidelity tracks for each speaker, and builds audiovisual datasets that synchronize speech, facial expression, gesture and timing. Its Data Foundry turns those recordings and domain-expert work into training data, alignment signals, evaluation suites and benchmarks. The company says its expert network includes thousands of vetted contributors and that its work is used by unnamed frontier AI labs and Fortune 100 enterprises; it also says revenue has reached seven figures. Those metrics are company-reported, and the customers were not identified in the announcement.
The clearest public evidence is the Converse benchmark family. Converse-STT compares 15 speech-to-text models on Ocular's two-person American English conversations and on public Pipecat audio. Ocular reported that 12 of the 15 models produced higher word-error rates on its conversational recordings. That does not establish that every production voice model will fail in the same way, and the dataset remains small, but it shows why clean or widely reused public clips may not predict performance when two people overlap, hesitate or correct themselves in real time. Cekura, Ocular's benchmark collaborator, separately described using an unseen dataset annotated by Ocular for the comparison.
Ocular started somewhere else entirely. Its original product focused on enterprise search and actions across workplace tools. CEO and co-founder Michael Moyo and CTO and co-founder Louis Murerwa began the company in Y Combinator's Winter 2024 batch. The current business reaches further down the AI stack, toward the datasets and evaluations labs use to train voice-native and audiovisual systems.
According to the announcement, Drive Capital's lead investment gives Ocular room to expand the Converse research program, develop more datasets and evaluations, and hire across engineering and research. Ocular said the capital also supports a wider research and evaluation program, more high-fidelity datasets and a growing team.
About the Company
Builds expert training data and benchmarks for voice and multimodal models; expanding expert-generated data, benchmarks, and evaluations for frontier AI.