Back
Oct 1, 2026

SSML Compatibility Across TTS Providers: Why Markup Breaks When You Switch Engines

TLDR: SSML compatibility is the gap between the W3C spec and what each TTS API actually honors. Amazon Polly, Azure Speech, Google Cloud TTS, and ElevenLabs all accept markup, but they do not share a tag matrix. Teams who paste one SSML document across vendors get silent no-ops: missing pauses, ignored phonemes, and vendor extensions that never fire.

SSML compatibility is the degree to which Speech Synthesis Markup Language written for one text-to-speech engine produces the same pauses, pronunciation, and prosody on another engine. The core answer is that compatibility is partial by design: the W3C SSML 1.1 recommendation defines the language, and each vendor ships a subset plus proprietary extensions. That gap matters because production pipelines treat markup as a portable quality control, then discover it is not.

This post is not another SSML tag tutorial. For tag mechanics, see our SSML for text to speech guide. Here the job is interoperability: what survives a provider switch, what fails silently, and how you keep pronunciation rules when the engine changes.

What does SSML compatibility actually mean?

SSML compatibility means two engines interpret the same markup document with equivalent timing, pronunciation, and emphasis. In practice, vendors document different supported elements, different phoneme alphabets, and different failure modes. The W3C recommendation itself says SSML specifies gross properties of synthetic speech such as pronunciation, volume, pitch, and rate, and that exact rendering still depends on the synthesizer.

Microsoft's Azure Speech transparency note still treats SSML as a best practice for output quality. That is true inside one engine. It is not a claim that markup is a cross-vendor contract.

Which SSML tags survive a provider switch?

A tag survives a provider switch only if both engines document it, implement it, and apply it on the model version you actually call. Use this matrix as a starting audit, then re-check vendor docs before every cutover.

ControlAmazon PollyAzure SpeechGoogle Cloud TTSElevenLabs
<break>SupportedSupportedSupportedSupported on most models; not on Eleven v3 or v4 (use audio tags / punctuation)
<phoneme>SupportedSupported, plus custom lexiconsSubset of W3C; do not assume parity with Polly/AzureFlash v2 and Turbo v2 (English); Eleven v3 uses native IPA
<say-as>SupportedSupportedSupported (dates, digits, cardinals, etc.)Not a documented production equivalent
<prosody>SupportedSupportedDocumented in Google's subsetPrefer model prompting over markup
<sub>SupportedSupportedSupportedNot a documented production equivalent
Vendor extensions<amazon:effect> and relatedStyle / role / lexicon URLsVoice-type limits inside <voice>Audio tags such as [pause] on v3

Google Cloud TTS states that not all W3C elements are supported, and SSML characters count toward quota. ElevenLabs help docs are explicit: all models except Eleven v3 support SSML break tags up to 3 seconds; v3 uses [pause], [short pause], and [long pause] instead. Eleven v4 and v3 do not support SSML break tags. Phoneme tags on ElevenLabs are limited to Flash v2 and Turbo v2 in English.

That last row is the lock-in vector. Polly's <amazon:effect> and Azure lexicon URLs from Blob Storage improve quality on that stack and become dead XML everywhere else.

Why do unsupported SSML tags fail silently?

Unsupported SSML tags fail silently because synthesizers treat unknown elements as ignorable markup rather than as request errors. The API still returns audio. Your pause never happens. Your IPA never binds. A 200 response is not proof the tag ran.

This is the production failure mode, not a 400. Teams comparing engines on the same SSML file score "ElevenLabs sounds rushed" when the real difference is that v3 dropped the break tag they still send. The same script with Polly-only extensions looks "flatter" on Google because those extensions never execute.

A compatibility QA block that actually catches this:

  1. Tag inventory. List every element and attribute in your SSML corpus, including namespaces (amazon:, Azure mstts:).
  2. Per-engine allowlist. Map each tag to the destination model's current docs, not last quarter's notes.
  3. Negative test. Render one clip with the tag stripped and one with the tag present. If waveforms and pause lengths match, the tag is a no-op on that engine.
  4. Model pin. Re-run the pair after a vendor model bump. ElevenLabs already split break behavior by version (v2 family vs v3 vs v4).

Microsoft documents phonemes and custom lexicons as first-class pronunciation tools on Azure. Google's SSML page is a subset list, not a promise of phoneme parity. Do not score engines against a shared SSML file and call the result a voice-quality ranking. You are often ranking markup support.

How do you keep pronunciation when engines disagree on SSML?

You keep pronunciation by storing rules outside any one vendor's markup dialect, then compiling them per engine. A portable glossary is the source of truth: term, preferred spoken form, IPA, locale, and forbidden expansions (for example, "Onepin" never read as "one pin").

Compile path by destination:

  • Polly / Azure: emit <phoneme> or lexicon entries those engines document.
  • Google Cloud TTS: prefer <sub> and <say-as> for the controls Google lists; do not ship Azure lexicon URLs.
  • ElevenLabs Flash/Turbo v2: emit English <phoneme> only on those models.
  • Eleven v3: emit IPA / audio tags, not SSML breaks.

Then validate the audio, not the XML. Markup raises the odds. It does not certify the clip. Pair this with the brand-name pronunciation workflow and treat a provider swap as a re-compile, not a credential change. For the full cutover sequence, use how to switch TTS providers.

Reddit-era advice to "just write better SSML" still assumes a Google-class engine. That advice fails the moment the same file hits a neural vendor that replaced SSML breaks with prompt tags.

What is the difference between an SSML file and a TTS production layer?

An SSML file is input markup for one synthesizer. A TTS production layer stores pronunciation, format, and quality rules above the model, routes each job to an engine that can honor those rules, and rejects clips that miss the reference even when the markup looked perfect.

Onepin is a voice workflow platform that orchestrates, validates, and ships production-ready audio across 100+ TTS models. You keep a provider-neutral glossary, compile engine-specific markup or prompts at render time, and score every output before it ships. Switching from Polly to ElevenLabs or Google is a routing plus compile step, not a rewrite of every script.

SSML remains useful. It is still the right way to express pauses and phonemes on engines that implement it. Treat it as a per-model dialect with a shared glossary behind it, then listen to the file that comes back. Start at onepin.ai or the docs.

Frequently asked questions

Does SSML work the same way on every TTS provider?
No. SSML is a W3C standard, but each vendor implements a subset. Amazon Polly and Azure Speech expose broad tag sets, Google Cloud TTS documents a subset of the spec, and ElevenLabs supports break and phoneme tags only on specific models. A tag that works on one engine is often ignored on another.
Which TTS models support SSML phoneme tags?
Amazon Polly and Azure Speech both support phoneme markup. ElevenLabs documents SSML phoneme tags on Flash v2 and Turbo v2 for English, and native IPA on Eleven v3. Google Cloud TTS documents a subset of W3C SSML and historically has not treated phoneme as a first-class production control the way Azure and Polly do.
What happens if I send an unsupported SSML tag?
Most engines ignore the unknown tag and keep synthesizing. The pause or pronunciation you expected never lands, and the API still returns 200. That is why markup cannot replace output validation when you route the same script across vendors.
Can I reuse Polly SSML on ElevenLabs or Google Cloud TTS?
Not safely. Polly extensions such as amazon:effect are vendor-specific. Eleven v3 and v4 do not support SSML break tags; they use audio tags and punctuation instead. Google counts SSML characters toward quota and supports only the tags listed in its docs. Rewrite or strip tags per engine, then re-score the audio.
How do I keep pronunciation rules when I switch TTS engines?
Keep a provider-neutral glossary of brand names, IPA, and number rules at the orchestration layer. Translate that glossary into the markup or prompting style each engine actually honors, then validate every clip. Do not treat a single SSML file as portable.

Ready to publish?

Turn any script into production-quality voice,
in any language, in minutes.

Run your first line