Omilia Raises $67M Series B: Scaling Voice AI Without Scaling Validation

Omilia just closed a $67 million Series B led by Expedition Growth Capital. The Athens- and New York-based voice AI company grew annual recurring revenue more than 10x since its Series A to over $60 million, without raising additional equity in the interim. Its customers include Capital One, Discover, RBC, Taco Bell, and PSEG.
The round funds global expansion, including Omilia's first U.S. office in the second half of 2026. A Tier 1 U.S. bank already handles over 1 million calls per day on the platform, and Omilia reports capacity for 50,000+ concurrent voice interactions with sub-second latency.
These are impressive platform numbers. They measure the agent. They do not measure the audio.
What Does Omilia's $67M Actually Scale?
Omilia's platform metrics tell you how many calls it can handle, how fast it responds, and how often callers stay inside the automated flow without reaching a human agent. Forrester named it a Leader in Conversational AI Platforms for Customer Service in Q2 2026. Gartner placed it as a Visionary in the Magic Quadrant in July 2026. The enterprise conversational AI market is projected to grow 192% from $17.05 billion in 2025 to $49.8 billion by 2031, according to CMSWire's market analysis.
Call containment, latency, and concurrent capacity are infrastructure metrics. They answer: did the system stay up, and did the caller stay in the flow? They do not answer: did the spoken audio correctly pronounce the caller's name, read back the right account number, or deliver the policy term without mangling it?
Why Do Platform Metrics Miss the Audio Output?
Platform metrics miss the audio output because they operate at the agent layer, not the audio layer. The agent decides what to say. The TTS model decides how it sounds. These are two separate systems with two separate failure surfaces.
Omilia launched Lexis, its proprietary generative TTS engine, in July 2026 with sub-45ms latency across 25+ locales. Building a proprietary voice model solves latency. It solves data sovereignty. It does not solve per-output correctness.
A model that generates audio in under 45 milliseconds can mispronounce a surname just as fast as one that takes 200 milliseconds. Speed describes how a model runs. It says nothing about whether the audio it returns is right. At 1 million calls per day for a single bank client, even a 0.1% pronunciation error rate means 1,000 calls daily where the voice says something wrong with no automated system catching it.
What Happens When You Scale Voice AI Without Output Validation?
Scaling without output validation turns a small problem into a large one. According to the Bureau of Labor Statistics, U.S. contact centers handle approximately 2.9 million agent positions, and enterprise voice AI platforms are absorbing an increasing share of that call volume. When Omilia reports 200+ enterprise deployments, that represents hundreds of distinct vocabularies: financial product names, insurance policy terms, medical terminology, restaurant menu items, utility account identifiers.
Each deployment is a separate pronunciation failure surface. Each locale among the 25+ Lexis supports is another. Taco Bell's menu vocabulary is different from Capital One's financial products. A system that pronounces "quesarito" correctly can still mangle "HELOC" or "Roth IRA." Platform-level QA checks that the call completed. It does not check that the audio said what it was supposed to say.
The Forrester Wave evaluates platform completeness, vision, and customer experience strategy. It does not include per-output pronunciation accuracy, model version locking, or audio format compliance in its scoring criteria. A Leader designation validates the platform. It does not validate the audio leaving the platform.
How Does a Production Layer Close the Gap?
A production layer sits between the model (whether proprietary like Lexis or third-party) and the caller. It scores every audio output against a locked reference before it reaches the listener.
Four components close the gap at enterprise scale:
-
Pronunciation validation per deployment vocabulary. Lock a pronunciation dictionary for each client's domain terms (account types, product names, proper nouns) and score every output against it. A financial services deployment and a restaurant deployment need different dictionaries.
-
Model version locking. Pin the validated model version per deployment. When a model update ships, re-validate before switching. Silent updates change pronunciation, pacing, and number formatting without triggering any alert in the platform metrics.
-
Per-output quality scoring. Score every clip, not a sample. At 1 million calls per day, sample-based QA covers a fraction of a percent. The errors that matter (a wrong account number, a mangled medication name) live in the tail.
-
Audio format compliance. Validate codec, sample rate, loudness normalization, and silence padding before delivery. Telephony infrastructure (G.711, 8kHz) has strict requirements that cloud TTS APIs do not enforce by default.
Onepin is a voice workflow platform that orchestrates, validates, and ships production-ready audio across 100+ TTS models. It handles pronunciation validation, model version locking, per-output scoring, and format compliance as a layer above whatever model the deployment uses, whether proprietary or third-party.
The $67M Question
Omilia's $67M validates that enterprise voice AI is real, scaled, and generating meaningful revenue. The platform works. The question is whether the audio leaving the platform is validated before the caller hears it.
$67 million buys more deployments, more locales, more concurrent capacity. It does not automatically buy per-output pronunciation accuracy, model version control, or audio format compliance. Those are production-layer problems, not platform-layer problems.
The platform measures the agent. The production layer measures the audio. Two different guarantees, both required at enterprise scale.
Learn more at onepin.ai.
Frequently asked questions
- What did Omilia announce in August 2026?
- Omilia announced a $67 million Series B funding round led by Expedition Growth Capital on August 6, 2026. The capital will fund global expansion, including the company's first U.S. office, and continued growth of its voice-first agentic CX platform serving over 200 enterprise deployments.
- Does Omilia validate the audio output of its voice AI platform?
- Omilia reports platform-level metrics like call containment rates, sub-second latency, and concurrent voice interaction capacity. These measure agent performance, not whether the spoken audio output correctly pronounces names, numbers, or policy terms on each individual call.
- How many calls does Omilia handle per day?
- Omilia reports that a single Tier 1 U.S. bank customer handles over 1 million calls per day on its platform. At that volume, even a fraction of a percent error rate in audio output means thousands of calls daily with potential pronunciation or number errors.
- What is the difference between platform validation and output validation in voice AI?
- Platform validation measures whether the agent system works correctly, including call routing, containment, latency, and uptime. Output validation measures whether each individual audio clip the system delivers is correct, checking pronunciation of names, accuracy of numbers, voice consistency, and format compliance. Most voice AI platforms report the first but not the second.
- What is a voice AI production layer?
- A voice AI production layer sits above the TTS model and the agent platform. It scores every audio output against a locked reference, catches pronunciation errors, locks model versions to prevent silent drift, validates audio format compliance, and regenerates only the clips that fail. Onepin is a voice workflow platform that orchestrates, validates, and ships production-ready audio across 100+ TTS models.