Back
Sep 29, 2026

TTS Evaluation Methodology: Why WER and MOS Are Not Enough

TTS evaluation methodology is the set of checks you run on generated speech before it ships. Word Error Rate and Mean Opinion Score are useful floors for intelligibility and naturalness, but they do not tell you whether a clip is correct for a tracking ID, a brand name, a locale date, or a quoted emotion. Production teams need a layered scorecard, not a single vendor number.

TLDR

A TTS model can win a MOS bake-off and still fail a tracking number. Word Error Rate is computed by synthesizing audio, transcribing it with ASR, and comparing that transcript to the source text. ITU-T P.800 defines Mean Opinion Score as listener judgment of overall quality. Neither metric scores pronunciation of proper nouns, alphanumeric IDs, or job-specific context. EmergentTTS-Eval (1,645 cases across six challenge types) exists because WER and MOS miss those failures. Onepin sits above models such as ElevenLabs and Deepgram Aura so you can route, validate, and retry per clip instead of trusting a leaderboard average.

What is TTS evaluation methodology?

TTS evaluation methodology is a layered process that scores generated speech against the job you actually ship, not against a generic naturalness average.

Start with an intelligibility floor (ASR-based WER). Add a naturalness or preference check if listeners will hear long-form content. Then add checks that WER cannot see: homographs, locale normalization, brand lexicons, alphanumeric codes, URLs, and emotion or question prosody. Last, treat every failed clip as a routing event: retry, swap model, or block ship.

That last step is the difference between a research paper and a production pipeline. Papers report means. Pipelines need a pass/fail per file.

Why is Word Error Rate the wrong single metric?

WER answers one question: after ASR, does the transcript match the text you sent? It does not answer whether the clip sounded right.

Cartesia's evaluation write-up is blunt: two clips can share the same WER while listeners strongly prefer one, because WER saturates once models are mostly intelligible. WER is also tied to the ASR you used. A weak recognizer on accents, names, or expressive speech penalizes TTS for ASR mistakes. A strong recognizer can "correct" unclear audio and hide a synthesis miss.

Towards Responsible Evaluation for Text-to-Speech (Yang et al., 2025) makes the same point on the research side: ASR error can mismatch metric scores even when humans find the speech adequate, and further WER drops at already-low rates have negligible perceptual impact. Their appendix example: moving from 1.61 to 1.47 barely changes what users hear.

Deepgram's alphanumeric benchmark is the production counterexample. Commercial ASR on structured alphanumeric sequences lands around 43 to 58 percent accuracy versus 95 to 99 percent on general speech, a 3 to 10x error jump (Deepgram, 2026). If your eval set is news-style prose, WER will look excellent while order IDs still fail.

LayerWhat it measuresWhat it misses
WER / CERIntelligibility via ASR transcriptEmotion, homographs, brand names, codes
MOS / CMOSListener naturalness (ITU-T P.800)Correctness of IDs, cost, per-clip fail
SIM / speaker embeddingVoice similarityLong-form drift, pronunciation
Named challenge suiteEmotions, foreign words, URLs, syntaxYour domain lexicon unless you add it
Production QASpec, retry, routingNothing if you actually run it

Why does Mean Opinion Score stall in production?

MOS is a 1 to 5 listener rating of overall quality, with terminology standardized in ITU-T P.800.1. It remains the default for "does this sound human?"

It stalls as a ship gate for three reasons.

First, it is expensive and noisy. EmergentTTS-Eval notes that traditional MOS is costly and statistically noisy because rater pools change. Second, MOS does not ask whether "07/08/2026" should be July 8 or 7 August. Third, predicted MOS models (UTMOS, DNSMOS, NISQA) are proxies trained on other audio. Cartesia flags domain mismatch: a predictor trained on clean English read speech is not a universal score for multilingual, conversational, or codec-compressed clips.

EmergentTTS-Eval's own table is the practical proof. Deepgram Aura-2 can post a high MOS while losing win-rate on expressiveness and complex pronunciation versus a prompted GPT-4o-mini-TTS baseline. Win rates and MOS measure different things. If you only keep MOS, you keep the wrong winner for IVR, games, or localization.

What should a production TTS evaluation stack include?

A production stack is a sequence of gates, each with a named fail reason.

1. Intelligibility floor. Run ASR WER on every clip. Fail hard on substitutions in the source script. Do not stop here.

2. Pronunciation and context pack. Homographs (wind the clock vs the wind), years, currency, "St.", locale dates. Cartesia lists these as cases where fluent audio is still wrong.

3. Alphanumeric and URL pack. Tracking numbers, SKUs, emails, formulas. EmergentTTS-Eval's Complex Pronunciation category (emails, phones, URLs, STEM, acronyms vs initialisms) exists because open-source models collapse here. Deepgram recommends transcription comparison on at least 500 production-like codes and a >98 percent pronunciation target on those sequences.

4. Expressiveness pack if the job needs it. Emotions, questions, paralinguistics. EmergentTTS-Eval covers six categories with 1,645 cases and uses a large audio language model as judge, reporting high correlation with human preference. Use that style of A/B, not a single MOS mean, when the product is narrative.

5. Locale and lexicon pack. Brand names, drug names, product SKUs. This is where a generic leaderboard is useless. Your glossary is the eval.

6. Fail, route, retry. The clip either matches spec or it does not. Swap Cartesia or another model, regenerate, or block. Onepin is built for that loop across 100+ TTS models so you are not locked to whichever engine won last quarter's MOS.

See the TTS quality validation checklist for the operational companion to this methodology, and how to switch TTS providers when a gate keeps failing on one vendor.

How do I evaluate TTS without locking to one model?

You keep the scorecard constant and treat models as interchangeable workers.

Vendor blogs will always optimize the metric they look best on. Aura-2 leans entity-aware alphanumerics. ElevenLabs leans multilingual naturalness. OpenAI's prompted 4o-mini-TTS was the only closed model in EmergentTTS-Eval to clear 50 percent win-rate on Complex Pronunciation. None of those facts tell you which engine should speak your SKU list tomorrow.

Hold the packs above as fixtures. Run the same 500+ IDs, the same homograph list, the same locale dates, every time a provider ships a silent model update. Score per clip. Route the fail. That is evaluation as production infrastructure, which is what a TTS platform is for, versus what a TTS model is for.

Ready to score audio against a spec instead of a demo? Start on onepin.ai.

Frequently asked questions

Is Word Error Rate enough to evaluate TTS quality?
No. WER only checks whether an ASR transcript matches the source text. Two clips can share the same WER and still fail on emotion, homographs, brand names, or alphanumeric IDs. Use WER as an intelligibility floor, then add pronunciation, context, and production checks.
What is Mean Opinion Score in TTS evaluation?
MOS is a listener rating of overall speech quality, defined in ITU-T P.800 terminology. It is useful for naturalness comparisons but expensive, noisy across rater pools, and blind to whether a tracking number or drug name was spoken correctly.
What should a production TTS evaluation stack include?
A production stack starts with ASR-based intelligibility, then adds pronunciation and homograph checks, alphanumeric and URL suites, locale-specific normalization, and a fail-retry loop. Named suites such as EmergentTTS-Eval cover emotions, paralinguistics, foreign words, syntax, questions, and complex pronunciation.
How is TTS evaluation different from ASR evaluation?
ASR evaluation asks whether speech was transcribed correctly. TTS evaluation asks whether generated speech is usable for a specific job. Low WER can hide wrong homographs, wrong locale dates, or codes that listeners cannot copy. Production teams score per clip against a spec, not against a leaderboard average.

Ready to publish?

Turn any script into production-quality voice,
in any language, in minutes.

Run your first line