TTS SSML Character Billing: Why Markup and Retries Inflate Your Invoice

TLDR: TTS character billing meters the request you send, not the words a listener hears. Google Cloud Text-to-Speech counts SSML tags toward quota. Amazon Polly prices Neural at $16 per million characters. ElevenLabs Flash/Turbo lists $40 per million. Every retry is a second bill.
TTS SSML character billing is the practice of charging for characters in the synthesis request, including markup on vendors that count tags. The core answer is that list rates look cheap until SSML padding and failed retries multiply the same script. That matters because finance sees a per-million-character SKU while production teams send XML, IPA, and three takes of the same line.
This is not another catalog of vendor SKUs. For a broader rate card, see the TTS API pricing guide. This post covers two invoice leaks: markup that never becomes audio, and retries that bill the same job twice.
What does SSML character billing actually count?
SSML character billing counts the characters the API receives, not the phonemes that play. Google Cloud Text-to-Speech pricing is explicit: the total number of characters in the input string counts, including spaces and newlines, and all SSML tags except the wrapping <speak> tag also count.
A 200-character narration wrapped in <prosody>, <break>, and <say-as> can double the billed length on Google before anyone hears a word. Chirp 3: HD is US$30 per 1 million characters after the first million. Neural2 is US$16 per million. WaveNet and Standard sit at US$4 per million. The SKU you pick multiplies the markup tax.
Amazon Polly bills characters converted to speech or Speech Marks. Polly pricing lists Standard at $4 per million characters, Neural at $16, Generative at $30, and Long-Form at $100, outside the free tier. Polly's example table prices 1 million characters of Neural at $16 whether you send it as 1,000 requests of 1,000 characters or 10,000 requests of 100. Request shape does not save you. Character volume does.
Azure Speech documents text-to-speech as billed per character. Public list dollars on the Azure pricing page often render behind a signed calculator, so do not invent a Neural HD rate. Treat Azure as character-metered and confirm the SKU in your contract.
How expensive is a million characters on Polly, Google, and ElevenLabs?
A million characters is the unit vendors quote; it is not a million spoken words. Use this snapshot from vendor pages retrieved for this post. Rates change. Pin the URL and date in your cost model.
| Engine | Documented unit | Rate used here | Source |
|---|---|---|---|
| Amazon Polly Neural | 1 million characters | $16 | AWS Polly pricing |
| Amazon Polly Standard | 1 million characters | $4 | Same page |
| Amazon Polly Generative | 1 million characters | $30 | Same page |
| Amazon Polly Long-Form | 1 million characters | $100 | Same page |
| Google Cloud TTS Chirp 3: HD | 1 million characters | $30 | Google TTS pricing |
| Google Neural2 | 1 million characters | $16 | Same page |
| Google WaveNet / Standard | 1 million characters | $4 | Same page |
| ElevenLabs Flash / Turbo | $0.04 per 1,000 characters | $40 per million | ElevenAPI pricing |
| ElevenLabs v3 / v2 Multilingual | $0.08 per 1,000 characters | $80 per million | Same page |
ElevenLabs also lists promotional v4 prices ($0.022 per 1,000 characters until Oct 12 on the page we retrieved). Do not forecast on a countdown discount. Use the uncrossed list rate unless procurement already locked the promo.
Worked example, Google Chirp 3: HD: a 2,000-character script with 2,000 characters of SSML (breaks, say-as, nested prosody) bills 4,000 characters per attempt, minus the wrapping speak tag. At $30 per million, that is $0.12 per take. Three failed takes before a pass: $0.36 for one line. Scale to 10,000 lines and the markup-plus-retry tax is the budget, not the voice SKU.
Polly's own examples put a typical news article at about 6,500 characters and $0.10 Neural. That assumes you send the article text. If localization injects SSML on Google instead, you are no longer in Polly's example.
Why do TTS retries cost as much as the first render?
TTS retries cost as much as the first render because APIs meter each request. There is no "same script" discount. A 400 from bad credentials is cheap compared with a 200 that returns audio you then reject for pronunciation.
That is the production pattern: the clip is valid HTTP and invalid product. You call the engine again with a phoneme hint. Google bills the second payload, tags included. Polly bills the second character count. ElevenLabs bills the second thousand-character block.
Pair this with SSML compatibility across TTS providers. Unsupported tags often fail silently. You pay for markup that never executed, then you pay again when you rewrite it for the engine that actually honors <phoneme> or [pause].
A cost-aware retry policy:
- Strip then compare. Render once with tags, once without. If duration matches, the tags are no-ops. Stop shipping them.
- Glossary first. Store brand names and IPA outside SSML. Compile the smallest dialect the destination documents. See how to switch TTS providers.
- Cap attempts. Two automatic retries, then human or a different engine. Do not loop Neural HD until the invoice notices.
- Cache passers. Polly states you can cache and replay generated speech at no additional synthesis cost. Replay the good take instead of resynthesizing the changelog.
How do you keep character cost down without locking into one TTS model?
You keep character cost down by treating markup as a compiled artifact and treating validation as the gate before a paid retry. List rate shopping (Neural $16 vs Chirp $30 vs Flash $40 per million) is step one. Step two is refusing to send a Polly-shaped SSML blob to Google, where every tag is a billable character.
Onepin is a voice workflow platform that orchestrates, validates, and ships production-ready audio across 100+ TTS models. You hold a provider-neutral glossary, compile engine-specific markup only when that engine documents the tag, score pronunciation before a second call, and route the job to a cheaper SKU when the line does not need Long-Form or Chirp 3.
Character billing will not go away. Google, Polly, and ElevenLabs all meter input. The leak is sending XML you do not need and paying again for clips you already know fail. Start at onepin.ai or the docs.
Frequently asked questions
- Do TTS APIs charge for SSML tags or only spoken text?
- Google Cloud Text-to-Speech counts every character in the input string, including spaces, newlines, and SSML tags, except the wrapping speak tag. Amazon Polly bills on characters converted to speech or Speech Marks. Markup that never becomes audible still hits the meter on Google.
- How much does Amazon Polly Neural TTS cost per million characters?
- Amazon Polly lists Neural voices at $16 per 1 million characters outside the free tier, Standard at $4, Generative at $30, and Long-Form at $100. GovCloud Neural is $19.20 per 1 million characters. Those rates apply to speech and Speech Marks requests.
- Does a failed TTS retry cost extra?
- Yes. Character billing is per request, not per unique script. If the first render fails pronunciation QA and you call the same engine again, you pay for both inputs. Google also counts SSML on each attempt, so a markup-heavy retry is more expensive than a plain-text retry.
- How do ElevenLabs TTS API rates compare per million characters?
- ElevenLabs documents Flash and Turbo at $0.04 per 1,000 characters, which is $40 per million. v3 and v2 Multilingual list at $0.08 per 1,000 characters, or $80 per million. Promotional v4 list prices change; always read the live API pricing page before you forecast.
- How do I cut TTS character cost without locking into one vendor?
- Keep pronunciation rules in a glossary, compile the smallest markup each engine actually honors, validate before you ship, and only retry on engines that failed a named check. An orchestration layer routes jobs so you do not paste one fat SSML file into every vendor.