TTS Provider Failover: How Production Voice Pipelines Survive an Outage

TLDR: TTS provider failover is an ordered retry path across vendors so a 5xx, timeout, or quota error does not stall synthesis. Google Cloud Text-to-Speech publishes a 99.9% monthly SLO; Amazon Polly sits under the same 99.9% language-SLA credit table. Those credits do not replay your IVR or course audio. This guide covers retryable errors, SLA math, and why failover without pronunciation QA still ships bad clips.
TTS provider failover is the process of retrying a synthesis request on a backup vendor when the primary returns a retryable failure. The core answer is simple: keep an ordered vendor list, retry only the errors that will not repeat, then score the backup clip before it ships. It matters because a 99.9% SLO still leaves about 43 minutes of counted downtime in a 30-day month, and quota or alpha voices often sit outside that count.
This is not a provider migration. Migration rebuilds voice IDs and dictionaries on purpose. Failover is the 2 a.m. path when ElevenLabs returns 503 and the lesson still has to render.
What is TTS provider failover?
TTS provider failover is an ordered retry across vendors for the same script when the primary call fails in a retryable way. The first hop stays on the preferred engine. The second hop is a pre-wired backup with a known sample rate, codec, and voice mapping. The job succeeds if any hop returns audio that then passes QA.
LLM stacks already treat this as infrastructure. OpenRouter documents model fallbacks that fire on downtime, rate limits, and moderation refusals, and bills the model that actually answered. Inworld's 2026 gateway comparison lists automatic failover on Vercel AI Gateway, Inworld Router, OpenRouter, Portkey, and LiteLLM. TTS production still often hard-codes one SDK. That is the gap.
On August 4, 2026, DeepSeek's status page logged two degraded-performance windows: 1 hour 18 minutes (02:02-03:20 UTC) and 36 minutes (03:43-04:20 UTC), as recorded by AIHubMix. The lesson transfers: one endpoint is one failure domain even when another host still serves a similar model.
How much downtime does a 99.9% TTS SLA actually cover?
A 99.9% monthly SLO still allows about 43 minutes of counted downtime in a 30-day month. Google Cloud Text-to-Speech states that SLO in its official SLA. Downtime there means a server-side error rate above 5%, measured in consecutive minutes. Intermittent errors under a minute do not start a downtime period. Repeated identical requests do not count unless the client follows exponential back-off from 1 second up to 32 seconds.
Credits, if you file within 30 days with project ID and timestamps:
| Monthly uptime | Credit on the covered TTS bill |
|---|---|
| 99% to under 99.9% | 10% |
| 95% to under 99% | 25% |
| Under 95% | 50% (cap) |
Amazon Polly is listed under the Amazon ML Language SLA (last updated November 28, 2023). Availability is averaged across 5-minute intervals. A 5xx on the Polly API is an Error. The credit table matches Google at 10% and 25%, then jumps to 100% under 95%. Credits apply to future bills. They do not regenerate the clip.
Google also excludes alpha and beta features, quota denials, and client-side invalid fields. ElevenLabs publishes a public status board and sells custom SLA language on enterprise pricing. Status green is not the same as a contractual SLO in your MSA.
What errors should trigger a TTS failover?
Retryable TTS errors are 5xx, timeouts, and documented capacity or quota responses. Do not fail over auth failures, unknown voice IDs, or SSML the primary already rejected. Those 4xx cases repeat on the backup and burn characters.
Use this matrix:
| Signal | Fail over? | Why |
|---|---|---|
| HTTP 500 / 503 / 429 with Retry-After | Yes | Primary capacity or internal error |
| Client timeout / empty body | Yes | You never received audio |
| 401 / 403 | No | Fix keys, not vendors |
| 400 invalid SSML or voice | No | The script is wrong |
| Audio returns, QA fails | Route, then QA | Different failure class |
TTS quality validation still applies after a successful backup call. A 200 from Amazon Polly on a script tuned for ElevenLabs can pass HTTP and fail the brand name. Failover without a pronunciation check just moves the incident into the delivery folder.
OpenRouter prices the model that finally answered. TTS character billing works the same way: you pay the vendor that produced the file. Track both attempts so finance does not treat retries as unexplained spend. See the TTS API pricing guide for how retries show up on invoices.
How do I wire failover without breaking voice consistency?
Wire failover as config, not as an incident script. Keep a primary voice ID, a backup vendor, a mapped backup voice, and a shared format contract (sample rate, codec, loudness). Store pronunciation dictionaries at the orchestration layer so they are not trapped in one vendor console.
A practical order for most product and e-learning stacks:
- Primary: the engine you already QA'd for the locale.
- Same-family backup if the vendor offers a second region or a stable older model.
- Cross-vendor backup with a pre-scored voice pair, not a random catalog pick.
- Dead-letter the job if every hop fails QA. Silence in the CMS is better than a wrong proper noun in the course.
LLM gateways return attempt chains in metadata so you can debug which hop won. TTS needs the same: vendor, model version, latency, and QA score on every clip. What TTS orchestration is is this control plane, not another synthesizer.
Onepin is a voice workflow platform that orchestrates, validates, and ships production-ready audio across 100+ TTS models. Failover is a routing event. Pronunciation dictionaries and format rules stay above the vendor. The backup clip still has to pass the same gate as the primary. You are not locked into one engine when that engine degrades.
If you already generate on a single SDK, add the backup mapping before the next incident, not during it. Docs live at onepin.ai/docs.
Frequently asked questions
- What is TTS provider failover?
- TTS provider failover is an ordered retry path that sends a synthesis request to a backup vendor when the primary returns a retryable failure such as a 5xx, timeout, or quota error. It keeps the same script moving instead of stalling the job. It is not the same as a planned provider migration, which rebuilds voice IDs and dictionaries.
- Does a 99.9 percent TTS SLA mean my voice pipeline stays up?
- No. Google Cloud Text-to-Speech states a 99.9 percent monthly uptime SLO, and Amazon Polly is covered by the Amazon ML Language SLA at the same 99.9 percent credit threshold. Those policies pay bill credits, not replacement audio. Quota errors and alpha or beta voices are often excluded, so your app still needs its own retry path.
- What is the difference between provider failover and model fallback?
- Provider failover retries the same model or voice family on another vendor or region. Model fallback switches to a different model after every eligible channel for the primary has failed. LLM gateways document both layers. For TTS, a model switch also changes timbre, so you still need a pronunciation and format check on the backup clip.
- Should I fail over TTS on every error code?
- No. Retry 5xx, timeouts, and documented quota or capacity errors. Do not blindly retry 4xx authentication, invalid voice IDs, or SSML the primary rejected. Those failures will repeat on the backup and waste characters. Log the status, then fix the request.
- How does Onepin handle TTS failover?
- Onepin is a voice workflow platform that orchestrates, validates, and ships production-ready audio across 100+ TTS models. It treats vendor errors as a routing event, then scores the backup clip for pronunciation and format before delivery. You are not locked to a single engine when one vendor degrades.