TTS Pronunciation Lexicons: Why One Glossary Does Not Travel Across Providers
TLDR: A TTS pronunciation lexicon is a reusable map from graphemes to phonemes or aliases, usually W3C Pronunciation Lexicon Specification (PLS) 1.0 (Recommendation, 14 October 2008). Amazon Polly stores region-scoped PLS files and applies up to five per SynthesizeSpeech call. Azure Speech points SSML at a public .xml or .pls URI. Google Cloud Text-to-Speech injects custom pronunciations in the RPC. ElevenLabs phoneme dictionary tags only fire on specific models. Keep the glossary above the vendor or failover ships the wrong brand name.
A TTS pronunciation lexicon is a reusable dictionary that tells a synthesizer how to say a written form before audio is generated. The core answer is that vendors all claim "custom pronunciation," then store the rules in incompatible APIs: Polly PutLexicon in a region, Azure blob URI, Google RPC list, ElevenLabs project dictionary. It matters because a 200 OK on a backup engine does not load the primary's lexicon, so the product name that was correct yesterday is wrong on the retry.
This is not an SSML tag tutorial and not a cross-provider SSML matrix. Those posts cover markup that lives inside one request. A lexicon is the shared glossary you expect to survive the next model swap.
What is a TTS pronunciation lexicon?
A TTS pronunciation lexicon is a file or API object that maps graphemes (the written tokens) to either a phonetic string or an alias expansion, then applies that map before synthesis. W3C PLS 1.0 is the interchange format Polly and Azure still name in docs. One lexeme is one word or acronym. One lexicon is a set of lexemes for a language.
Polly's own examples are the production cases: stylized spellings such as "g3t sm4rt" aliased to "get smart," and "W3C" expanded to "World Wide Web Consortium." Humans parse those. Engines read the characters. The lexicon is the contract.
Inline <phoneme> in one script is a one-off. A lexicon is the list legal, localization, and brand teams think they already maintain.
How do Polly, Azure, Google, and ElevenLabs store lexicons?
They do not share a store. Each vendor attaches the glossary to its own region, URI, RPC, or project.
| Vendor | Store | Apply | Phonetic alphabets (docs) | Gotcha |
|---|---|---|---|---|
| Amazon Polly | PutLexicon in one AWS region | Up to five lexicon names on SynthesizeSpeech; first listed wins on the same grapheme | PLS with IPA or alias | A lexicon in us-east-1 is invisible in eu-west-1 |
| Azure Speech | Public .xml / .pls URI (Blob + SAS common) | SSML <lexicon uri="..."> | IPA, SAPI, UPS, x-sampa on <phoneme> | lexicon is not supported on the Long Audio API |
| Google Cloud TTS | custom_pronunciations on the synthesis RPC | Engine rewrites text to <phoneme> | IPA and X-SAMPA | Dictionary is per request, not a named regional object |
| ElevenLabs | Project / API pronunciation dictionary (TXT or PLS) | Auto-apply on the project, or locator on the API | IPA and CMU Arpabet | Phoneme dictionary tags skip on models other than eleven_v4, eleven_flash_v2, and eleven_v3; use alias there |
Polly documents order explicitly: two lexicons both defining Bob as Robert vs Bobby produce different audio depending on --lexicon-names LexA LexB vs the reverse. That is not a bug. It is first-match precedence.
ElevenLabs is explicit that Multilingual v2-class models skip dictionary phoneme tags and fall back to default pronunciation. Alias substitution is the documented workaround. A PLS file that "works in the ElevenLabs project" can be a no-op on the model you actually call from an agent.
Why does a lexicon fail when I switch TTS providers?
A lexicon fails on switch because the backup vendor never received the primary's object. Polly applies names that exist in that region. Azure fetches a URI. Google wants the pronunciations array on this RPC. ElevenLabs binds a dictionary ID to a project and then gates phonemes by model.
How to switch TTS providers already covers voice IDs. Pronunciation is the second lock-in: the glossary lives in a console your failover path does not query.
Failover without a compiled glossary is a quality incident, not an availability one. The backup returns audio. QA that only checks HTTP or WER still ships "Onepin" as three English words, or a drug name with the wrong stress. TTS evaluation methodology is the scorecard. The lexicon is the input that scorecard assumes you already ported.
Practical compile rules:
- Keep a vendor-neutral source of truth: grapheme, locale, preferred IPA, alias fallback.
- Emit Polly PLS per region; do not assume replication.
- Host Azure
.plswhere the speech resource can GET it (SAS if it cannot be public). - Attach Google
custom_pronunciationson every request, including retries. - For ElevenLabs, store both phoneme and alias rows, then select by model.
How should production teams own the glossary?
Own the glossary in the orchestration layer, then compile. Do not treat the Polly console as the system of record.
A working production loop:
- Brand and medical terms enter one list with locale and owner.
- CI emits per-vendor artifacts and fails the build if IPA is empty for a locale that requires phonemes.
- Synthesis always passes the compiled artifact, including backup hops.
- A pronunciation check runs on the clip, not only on the XML.
Onepin is a voice workflow platform that orchestrates, validates, and ships production-ready audio across 100+ TTS models. Dictionaries sit above PutLexicon, blob URIs, and project IDs. Failover still has to pass the same brand-name gate. You are not locked to one lexicon API when the primary engine degrades.
If the glossary still lives in a single vendor console, export it before the next incident, not during it. Docs: onepin.ai/docs.
Frequently asked questions
- What is a TTS pronunciation lexicon?
- A TTS pronunciation lexicon is a reusable map from written forms to how a synthesizer should say them, usually as W3C PLS XML with graphemes plus phonemes or aliases. It is applied before synthesis so brand names and acronyms stay consistent. It is not the same as one inline phoneme tag in a single script.
- Do Amazon Polly, Azure, Google Cloud TTS, and ElevenLabs share one lexicon file?
- No. Polly stores region-scoped PLS lexicons and applies up to five per SynthesizeSpeech call with first-match order. Azure references a public .xml or .pls URI from SSML. Google injects custom pronunciations in the RPC so text becomes phoneme tags. ElevenLabs dictionaries use phoneme or alias rules and phonemes only apply on specific models such as eleven_v4, eleven_flash_v2, and eleven_v3.
- What is the difference between a phoneme rule and an alias?
- A phoneme rule keeps the written word and supplies IPA, CMU Arpabet, or another alphabet so the engine says that spelling correctly. An alias replaces the grapheme with other words, such as expanding W3C to World Wide Web Consortium. Polly documents both. ElevenLabs tells you to use aliases on models that skip dictionary phoneme tags.
- Why do pronunciation dictionaries break during TTS failover?
- Vendor lexicons live in that vendor's region, blob URI, or project ID. A backup engine does not load the primary's PutLexicon object. If you fail over without compiling the same graphemes into the backup format, the HTTP call can succeed while the brand name is wrong.
- How does Onepin handle pronunciation lexicons across TTS models?
- Onepin is a voice workflow platform that orchestrates, validates, and ships production-ready audio across 100+ TTS models. It keeps the glossary above any single console, compiles per-vendor rules, and scores the clip before delivery so you are not locked to one lexicon API.