Streaming TTS Latency: Time to First Byte vs Time to First Audio

TLDR: Streaming TTS latency is not one number. Time to first byte (TTFB) marks the first response bytes; time to first audio (TTFA) marks the first playable samples after container headers. Deepgram documents total latency as network + TTFB + synthesis. ElevenLabs Flash cites ~75ms inference; Inworld publishes 25ms and sub-100ms server-side P99 TTFB excluding network. This guide covers how to measure TTFA, when REST vs HTTP stream vs WebSocket fits, and why routing alone does not fix bad first chunks.
Streaming TTS latency is the delay from a synthesis request until the caller hears usable speech. The core answer is to stop treating a single vendor “latency” slide as the whole stack: split network, first byte, first real sample, and full synthesis, then pick a transport that matches how text arrives. It matters because natural turn gaps sit near 200ms in conversation research Gradium cites from Stivers et al., so every wasted millisecond on headers, reconnects, or wait-for-file clients burns the agent budget.
This is not a provider failover guide. Failover answers 5xx and timeouts. Latency answers whether the first word lands before the user talks over you.
What is streaming TTS latency?
Streaming TTS latency is the end-to-end delay from request start to playable audio on the client, with progressive chunks instead of a single finished file. Deepgram’s latency model is explicit: total_latency = network + ttfb + audio_synthesis. You measure each term, or you optimize the wrong one.
In one Deepgram walkthrough, a full request clocked 745ms total latency, with first-byte activity after 616ms and roughly 277ms first-byte latency after the SSL handshake portion of that trace. Those numbers are illustrative of method, not a promise for your region. The method is the product: break the path before you blame the model.
Batch pipelines care about full-file time. Voice agents care about first playable sample. If your client buffers the whole response, you paid for streaming and still ship REST behavior.
What is the difference between TTFB and time to first audio?
TTFB is the time to the first response byte; TTFA is the time to the first byte that is actually audio samples after container metadata. Gradium’s Time to First Audio post is blunt: streaming APIs often send WAV headers, Ogg identification pages, or MP3 ID3 tags first. A naive TTFB can read 50ms of headers while the first samples land 200ms later. The user experiences the later number.
Measure TTFA by parsing past headers before you start the timer stop:
| Format | Skip before “first audio” |
|---|---|
| WAV | Initial 44-byte header (typical PCM WAV) |
| Ogg/Opus | Identification and comment pages |
| MP3 | ID3 tags, then first valid MPEG frame |
Keep input text, sample rate, and codec constant across vendors. Gradium’s own Paris benchmark (websocket where available, 100 queries, first 5 discarded) reported P50 TTFA of 258ms for Gradium, 304ms for Eleven Turbo v2.5, 324ms for Eleven Flash v2.5, 420ms for OpenAI GPT-4o Mini TTS path, and 706ms for Eleven Multilingual v2. That table is one lab path. Use it as a measurement template, not as a global ranking.
How do vendor latency claims map to real clients?
Vendor latency claims usually describe server-side inference or on-server TTFB and exclude your network hop. Read the footnote before you set an SLO.
| Source | Claim (as published) | Scope note |
|---|---|---|
| ElevenLabs Flash | ~75ms | Model inference only; end-to-end varies by location and endpoint |
| ElevenLabs Flash + WebSockets (docs table) | 100-150ms TTFB in NA/EU/SEA; 150-200ms in parts of Asia | Geographic proximity table in the same latency guide |
| Inworld Realtime TTS-2 Flash | 25ms TTFB | P99 server-side; excludes network |
| Inworld Realtime TTS-2 | <100ms TTFB | Same server-side P99 definition |
| Cartesia Sonic | Sub-90ms latency (marketing) | Product page claim; validate with your own TTFA harness |
| Deepgram latency docs | Linear character growth ~40ms per 100 characters after a constant baseline | Character length effect on total latency, not a single TTFB number |
ElevenLabs also ranks voice type for speed: default, synthetic, and instant clones tend to outrun professional voice clones, and higher-fidelity output formats can add delay. Deepgram notes sample rate choices of 8-48 kHz do not by themselves change synthesis speed on their stack, while compressed formats can shrink transfer time on constrained links (REST only for some compressed paths).
If your chart says 75ms and your dashboard says 400ms, you are not “bad at APIs.” You are measuring a different object.
When should I use REST, HTTP streaming, or WebSocket for TTS?
Pick the transport from how text arrives and whether you need mid-utterance control, not from the word “realtime” on a homepage. Deepgram’s WebSocket vs REST guide (Mar 27, 2026) frames it cleanly: REST waits for a complete asset; HTTP chunked streaming starts progressive playback without bidirectional control; WebSocket keeps a persistent duplex channel for incremental LLM tokens, cancel, and flush.
| Protocol | Best when | Cost you accept |
|---|---|---|
| REST | Full script ready; you want a cacheable file and simple retries | User waits for full synthesis before first sound if you do not stream the body |
| HTTP chunked stream | Progressive play, no mid-turn cancel | Less control than WebSocket |
| WebSocket | Token-by-token LLM text; barge-in; multi-turn without reconnect | Connection lifecycle, multiplexing, session hygiene |
Deepgram cites 50-100ms per-request savings for WebSocket vs REST in multi-turn conversations when reconnect cost would otherwise stack. Gradium measures on the order of ~50ms just to open a fresh websocket per turn, and shows lower TTFA when connection setup is excluded and sessions multiplex on one socket. ElevenLabs documents three endpoint styles (regular, streaming SSE, websockets) and warns that with auto_mode off, a chunk schedule can stall generation until enough characters arrive.
For PSTN and contact center paths, pacing and interruption control often beat raw transport milliseconds. Protocol choice still matters; it is not the whole call quality story.
How do I cut streaming TTS latency without wrecking quality?
Cut latency with measurement, colocation, streaming clients, and smart chunking, then keep a pronunciation gate on the first playable chunk. Deepgram’s playbook: stream from first byte, self-host or place servers near users when network dominates, pack long content near the 2000-character input ceiling when you must chunk, and avoid splitting one sentence across requests (prosody breaks). Gradium adds: colocate STT, LLM, and TTS, stream LLM tokens into TTS, and multiplex turns on a warm socket.
A practical production checklist:
- Log TTFA, not only TTFB, with format-aware parsers.
- Log region, model ID, voice ID, and connection reuse.
- Prefer flash or realtime model SKUs for agent turns; keep higher-quality models for offline batch where total file time wins.
- Chunk on sentence boundaries for long replies; never mid-clause.
- Fail closed on empty first audio, extreme silence pads, or brand-name misses even when TTFB is green.
- Treat vendor P99 marketing as a ceiling for the server hop, not your customer SLO.
Onepin is a voice workflow platform that orchestrates, validates, and ships production-ready audio across 100+ TTS models. Latency routing picks a fast path for live turns. Pronunciation lexicons and format contracts stay above any single stream API. A 25ms first chunk that misreads the product name is still a failed turn. You are not locked to one engine when another region or model clears TTFA with the same QA score.
If your stack still waits for Content-Length before play(), fix the client before you swap vendors. Docs: onepin.ai/docs. Related reading: what TTS orchestration is, TTS quality validation checklist, and how to switch TTS providers.
Frequently asked questions
- What is time to first byte in streaming TTS?
- Time to first byte (TTFB) is the elapsed time from starting a TTS API request until the client receives the first byte of the response. Deepgram defines it that way in its latency docs. Those first bytes are often container headers, so TTFB alone does not always mean the user can hear speech yet.
- What is the difference between TTFB and time to first audio?
- TTFB timestamps the first response bytes. Time to first audio (TTFA) timestamps the first chunk that actually contains playable samples after you skip WAV headers, Ogg pages, or ID3 tags. Gradium notes a server can return metadata in about 50ms while real samples arrive hundreds of milliseconds later, so naive TTFB can understate what the user hears.
- When should I use WebSocket instead of REST for TTS?
- Use REST when the full script is ready and you need a complete file for caching or retries. Use HTTP chunked streaming when you want progressive playback without bidirectional control. Use WebSocket when LLM tokens arrive mid-turn and you need cancel, flush, or turn control on one open connection. Deepgram's 2026 protocol guide puts that decision on text arrival pattern, not marketing labels.
- Why is my measured TTS latency higher than the vendor chart?
- Most published figures are server-side inference or on-server TTFB and exclude your network hop. ElevenLabs states Flash ~75ms is inference only. Inworld labels 25ms and sub-100ms as P99 server-side TTFB excluding network. Add TLS, region distance, cold connections, and any wait-for-full-file client code.
- How does Onepin help with streaming TTS latency?
- Onepin is a voice workflow platform that orchestrates, validates, and ships production-ready audio across 100+ TTS models. It can route latency-sensitive turns to fast engines, keep pronunciation and format gates on every hop, and avoid locking you to a single vendor's stream path. Fast first audio still has to pass the same quality bar as batch output.