Back
Oct 5, 2026

TTS Language Routing: Why One Model Per Locale Beats a Global Default

TLDR: TTS language routing is the production rule that selects a model and voice per locale instead of one global engine. ElevenLabs language counts change by model: Multilingual v2 is 29, Flash v2.5 is 32, v3 is 70+, v4 is 90+. Cartesia Sonic 3.6 lists 44 languages. Amazon Polly binds each voice to a language code. A catalog count is not a production score.

TTS language routing is the decision layer that maps each utterance's locale to a specific engine, voice ID, and language code before synthesis. The core answer is that multilingual marketing copy hides per-model tables, so a single default engine will accept Korean or Odia text and still ship English-biased audio. It matters because HTTP success does not mean the locale was in that model's list.

This is not TTS provider failover (retry on error) and not a pronunciation lexicon (grapheme maps). Routing is the choice of which engine you even call.

What is TTS language routing?

TTS language routing is a lookup: locale in, model plus voice plus language parameter out. You run it before the first byte of audio. It is not "the vendor supports many languages," which is a catalog claim. It is "this utterance uses this engine because we measured it for this locale."

A useful mental model is three layers:

  1. Catalog: the vendor language list for that model ID.
  2. Binding: the voice ID and language / locale field the API requires.
  3. Score: a pronunciation check on the clip, not WER on a transcript of English.

Skip layer 3 and you will keep the cheapest global default forever.

How do vendor language tables actually differ?

They differ by model ID, not by brand homepage.

Vendor / modelDocumented language scopeBinding you must sendProduction gotcha
ElevenLabs Multilingual v229 languagesVoice from the library; text in that languageDocs list English, Japanese, Chinese, German, Hindi, French, Korean, Portuguese, Italian, Spanish, Indonesian, Dutch, Turkish, Filipino, Polish, Swedish, Bulgarian, Romanian, Arabic, Czech, Greek, Finnish, Croatian, Malay, Slovak, Danish, Tamil, Ukrainian, Russian
ElevenLabs Flash v2.532 languages (~75ms)Same, plus Hungarian, Norwegian, VietnameseThe "32 languages" voice-library line is not v2
ElevenLabs v3 / v4v3 70+, v4 90+Model ID must match the count you assumedA v2 call will not inherit the v4 table
Cartesia Sonic 3.644 languageslanguage or locale, never bothOdia (or) and Urdu (ur) are new vs 3.5; snapshot sonic-3.6-2026-08-27 pins the list
Amazon PollyPer-voice language codes (e.g. arb, ar-AE, yue-CN, en-IN)Voice ID that exists for that codeNeural vs Standard vs Generative is per row, not per brand
Google Cloud Chirp 3 HD30 distinct styles across many languagesVoice name plus language codeStyle count is not a locale count

ElevenLabs is explicit: pick a voice whose accent matches the target language. Cartesia is explicit: set language such as en or a regional locale such as en-GB. Polly will reject or degrade if you pair Joanna (en-US) with Hindi text. Those are routing inputs, not post-production fixes.

Why does a global default TTS model fail in production?

A global default fails because the API will often still return audio. The miss is phonetic, not HTTP.

Typical failure modes:

  • Wrong table: Flash v2.5 covers Vietnamese. Multilingual v2 does not list it. Same ElevenLabs account, different model ID.
  • Wrong binding: Cartesia gets text in Urdu without language: ur.
  • Wrong voice: An English-trained clone speaking Japanese script with English timing.
  • Wrong engine class: Polly Standard for a locale that only has Neural, or Neural for a locale that only has Standard (Icelandic is-IS is Standard-only on the public table).

TTS evaluation methodology already warns that WER on ASR of the clip measures lexical match, not native pronunciation. Language routing is the control that decides which model even enters that scorecard.

Failover without routing copies the same bad default onto the backup. Routing without failover still dies when the chosen region's quota hits. You need both.

How should production teams route TTS per locale?

Keep a locale table above any vendor console, then compile.

A working loop:

  1. Inventory locales you actually ship (BCP-47), not languages marketing listed.
  2. For each locale, store primary model ID, voice ID, language/locale field, and a scored backup that is a different vendor, not a sibling model with the same gap.
  3. Reject synthesis if the locale is missing from that model's documented list.
  4. Score the clip for pronunciation of brand terms in that locale (lexicon compile still applies).
  5. Pin snapshots when the vendor versions the table (Cartesia dated sonic-3.6-YYYY-MM-DD).

Onepin is a voice workflow platform that orchestrates, validates, and ships production-ready audio across 100+ TTS models. Locale routing sits above ElevenLabs model IDs, Cartesia language fields, and Polly voice codes. You are not locked to one language table when a new market goes live.

If every locale still hits the same model_id, you do not have routing. You have a default. Start at onepin.ai/docs.

Frequently asked questions

What is TTS language routing?
TTS language routing is the production rule that selects a model, voice, and locale code for each utterance instead of sending every language through one default engine. Vendors publish different language lists per model family. Routing matters because a 200 OK on the wrong model still ships English-biased phonetics or a missing locale.
Does ElevenLabs support the same languages on every model?
No. ElevenLabs documents Multilingual v2 at 29 languages, Flash v2.5 at 32 languages (adds Hungarian, Norwegian, and Vietnamese), Eleven v3 at 70-plus, and Eleven v4 at 90-plus. Marketing copy that says 32 languages often refers to the voice library, not the model you actually call.
How is language routing different from TTS failover?
Failover retries the same request on a backup when the primary errors or times out. Language routing chooses the primary before the call, based on locale. You still need failover after routing. A Korean line that always hits an English-first model is a routing miss, not an availability incident.
Why can a multilingual model still sound wrong in a second language?
A model can accept the script and still carry the default voice phonetic bias. ElevenLabs docs tell you to pick a voice whose accent matches the target language. Cartesia requires a language or locale field. Polly binds each voice ID to a language code. Passing text without matching the voice and code is the usual failure.
How does Onepin handle per-locale TTS routing?
Onepin is a voice workflow platform that orchestrates, validates, and ships production-ready audio across 100+ TTS models. It routes by locale, then scores pronunciation on the clip so you are not locked to one vendor language table when a market goes live.

Ready to publish?

Turn any script into production-quality voice,
in any language, in minutes.

Run your first line