Back
Oct 7, 2026

TTS Audio Format Guide: Sample Rate, PCM, WAV, MP3, and Loudness

TLDR: TTS audio format is the codec, container, sample rate, and bit depth your provider returns, plus the loudness target you ship. Defaults differ: OpenAI ships MP3 by default and recommends WAV or PCM for lowest decode delay; ElevenLabs defaults to MP3 with PCM, μ-law, A-law, and Opus options; Google Cloud Gemini-TTS documents LINEAR16 and related encodings on Cloud TTS, while Vertex can return PCM 16-bit 24 kHz without WAV headers. Lock a delivery contract before you rank voices.

TTS audio format is the production contract between a text-to-speech API and whatever plays or edits the file next. The core answer is to stop treating each vendor default as finished audio and lock codec, sample rate, container, and loudness to the channel you ship. It matters because a 24 kHz PCM stream with no header, an 8 kHz μ-law clip on a 48 kHz timeline, and an MP3 master that fails podcast loudness all get blamed on the model when the pipeline failed.

This is not a streaming latency guide. Latency answers first playable sample. Format answers whether that sample is decodable, editable, and legal for the destination.

What is a TTS audio format contract?

A TTS audio format contract is the exact codec, sample rate, bit depth, channel count, container, and loudness target you require from synthesis through publish. Without it, every provider ships a marketing default and every editor rewrites the file differently.

FieldWhat to lockFailure mode
CodecPCM/LINEAR16, MP3, Opus, μ-law, A-lawBad decode path
Sample rate8 / 16 / 22.05 / 24 / 44.1 / 48 kHzPitch or speed errors
ContainerRaw PCM vs WAV vs MP3Header mis-reads
ChannelsMono vs stereoTimeline and loudness math
LoudnessIntegrated LUFS + true peakChannel rejection

OpenAI’s Speech API documents MP3 as the default, with opus, aac, flac, wav, and pcm available, and states that WAV or PCM give the fastest response times because you avoid decode overhead. PCM there is raw 24 kHz, 16-bit signed, little-endian samples without a header. Pass that to a player that expects WAV and you get silence or static, not a weak voice.

How do major TTS providers differ on output formats?

Major TTS providers differ on defaults, tier-gated quality, and whether raw PCM includes a WAV header. Read each API’s format string.

ProviderDefault pathDocs facts
OpenAI SpeechMP3WAV/PCM for low latency; PCM = 24 kHz s16le, no header
ElevenLabs TTSMP3MP3 22.05-44.1 kHz (32-192 kbps); PCM S16LE; μ-law/A-law at 8 kHz; Opus at 48 kHz; higher quality on paid tiers
Google Cloud Gemini-TTSLINEAR16 on Cloud TTSUnary: LINEAR16, ALAW, MULAW, MP3, OGG_OPUS, PCM; streaming favors PCM/ALAW/MULAW/OGG_OPUS; Vertex can return 24 kHz PCM without WAV headers
xAI Grok TTSMP3 when omittedDefault MP3 at 24 kHz / 128 kbps

ElevenLabs strings look like mp3_22050_32 (codec_sample_rate_bitrate). Google notes rates such as 44,100 Hz for CD-class fidelity and MULAW for telephony-style paths. Vertex’s no-header PCM note is the classic trap: bytes are valid only if your decoder knows rate, width, and endianness. Log format fields on every render ID.

When should I use PCM, WAV, MP3, or telephony codecs?

Use PCM or WAV when you still need to process audio. Use compressed codecs when you only need delivery. Telephony codecs are a third lane for PSTN and contact-center paths.

PCM / LINEAR16 / WAV fit normalization, trim, concatenation, and archives (raw PCM needs an explicit rate). OpenAI calls out WAV or PCM when decode cost matters. MP3 fits finished delivery and CDN cache; re-encoding stacks artifacts. Opus / OGG_OPUS fit modern streaming (ElevenLabs lists Opus at 48 kHz). μ-law / A-law (8 kHz) fit telephony and IVR. If editors or QA will touch the file again, stay lossless until the last hop.

How should I handle sample rate and resampling?

Handle sample rate as a destination requirement first, then a provider option, then one controlled resample. Casual multi-hop resamples are how brand reads turn chipmunk or muddy.

DestinationCommon rateNotes
PSTN / many IVRs8 kHzOften with μ-law/A-law
Speech ML / STT loops16 kHzWideband workhorse
Many neural TTS APIs22.05 / 24 kHzOpenAI PCM is 24 kHz; Google Vertex examples use 24 kHz
Some master delivery44.1 kHzElevenLabs options include 44.1 kHz (tier limits may apply)
Video / film post48 kHzMatch the timeline

Deepgram’s latency docs note that sample rate choices from 8-48 kHz do not by themselves change synthesis speed on that stack, while compressed formats can shrink transfer time on constrained links. Speed is not the only scoreboard. A wrong rate with a correct codec still fails when the NLE assumes 48 kHz.

Request the closest native rate, resample once with fixed settings, and never chain MP3 → WAV → MP3.

What loudness gate should production AI voice use?

Production AI voice should use an integrated loudness target plus a true-peak ceiling per channel, measured the same way every time. Peak-normalized “make it loud” is not a standard.

Industry practice clusters around EBU R128 style loudness units (LUFS/LKFS). Youlean’s comparison table indexes broadcast and platform targets: TV-oriented EBU R128 work centers near -23 LUFS integrated, while many podcast and app voice workflows master hotter (commonly near -16 LUFS integrated). Streaming music platforms often normalize playback toward roughly -14 LUFS listening targets, a different job than speech IVR.

  1. Measure integrated loudness (LUFS) on the full clip.
  2. Measure true peak (dBTP), not only sample peak.
  3. Fail if integrated loudness sits outside your channel window.
  4. Fail if true peak exceeds your ceiling (lock one number, often near −1 to −2 dBTP in broadcast-adjacent specs).
  5. Compare same-script renders across providers only after the same loudness gate, or you crown the loudest model, not the best one.

Loudness is a validation-layer problem. Model ranking without a loudness contract is preference theater.

How do I keep multi-provider TTS delivery consistent?

Keep multi-provider delivery consistent with a shared format profile, a post-synthesis normalize step, and fail-closed QA before the asset is cacheable. Provider A’s default MP3 at 22.05 kHz and provider B’s raw 24 kHz PCM will not match in an editor unless you own the last mile.

Define named profiles (agent_realtime_pcm_24k, ivr_ulaw_8k, course_wav_48k_mono, social_mp3_44k_128), map each to exact API format enums, run format sniff plus loudness and pronunciation gates, then write to CDN only after gates pass.

Onepin is a voice workflow platform that orchestrates, validates, and ships production-ready audio across 100+ TTS models. Format and loudness contracts sit above any single engine, so a failover to a second model still has to pass the same delivery profile. You are not locked into one vendor’s default MP3 when the channel needs 48 kHz WAV masters or 8 kHz telephony frames.

If your quality tickets are really decode errors, rate mismatches, or loudness swings, fix the contract before you swap models. Docs: onepin.ai/docs. Related reading: streaming TTS latency (TTFB vs TTFA), TTS provider failover, SSML compatibility across providers, and what TTS orchestration is.

Frequently asked questions

What audio format should I request from a TTS API?
Match the destination, not the vendor default. Use PCM or WAV when you will resample, loudness-normalize, or stitch clips. Use MP3 or Opus for delivery and bandwidth-limited playback. Telephony often needs 8 kHz mu-law or A-law. Default MP3 is fine for demos and wrong for many production post paths.
What is the difference between PCM and WAV from a TTS provider?
PCM is raw samples without a container header. WAV wraps those samples with a header that tells players the sample rate, bit depth, and channel count. OpenAI documents PCM as 24 kHz 16-bit little-endian without a header. Google Vertex Gemini TTS can return PCM 16-bit 24 kHz without WAV headers, so you must wrap or decode with the correct rate.
Does sample rate change TTS synthesis speed?
On some stacks, sample rate is mostly a delivery and quality choice rather than a synthesis-time knob. Deepgram notes that 8 to 48 kHz choices do not by themselves change synthesis speed on its stack, while compressed formats can shrink transfer time. Always re-check the provider you use and measure with the same format you ship.
What loudness target should AI voice assets hit?
Broadcast speech often tracks EBU R128-style integrated loudness near -23 LUFS for TV. Podcast and app voice masters commonly sit louder, around -16 LUFS integrated in many podcast workflows. Pick one contract per channel, measure integrated loudness and true peak, and fail clips that miss the gate before publish.
How does Onepin help with TTS audio formats?
Onepin is a voice workflow platform that orchestrates, validates, and ships production-ready audio across 100+ TTS models. It can enforce format, sample rate, and loudness contracts above any single vendor default so you are not locked into one engine when another model passes the same delivery gate.

Ready to publish?

Turn any script into production-quality voice,
in any language, in minutes.

Run your first line