Models

32 featured voice models across 19 providers. A selection of what Onepin routes to, planned and validated through the same pipeline.

Price

32 models

Alibaba Cloud

Qwen-Audio-3.0-TTS Flash

1,500credits
/ 1M chars

Alibaba's real-time Qwen-Audio 3.0 tier: first audio in about 300ms across 16 languages, directed in plain language ('slower, warmer') or with inline tags like [giggles].

Voices
609 voices
Languages
3 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Plain-language instruction (emotion, role, pace), rate, pitch and volume, plus inline emotion tags.

Alibaba Cloud

Qwen-Audio-3.0-TTS Plus

2,000credits
/ 1M chars

The quality tier of Qwen-Audio 3.0, tuned for naturalness and timbre fidelity over speed. Same 16 languages and plain-language direction as Flash, and #1 on the Artificial Analysis TTS leaderboard at launch.

Voices
598 voices
Languages
3 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Plain-language instruction (emotion, role, pace), rate, pitch and volume, plus inline emotion tags.

AWS Polly

Polly Long-Form

10,000credits
/ 1M chars

Amazon's long-form TTS for audiobooks and podcasts, with custom lexicons and full SSML phoneme control across 30+ languages.

Voices
6 voices
Languages
2 regions supported
Custom pronunciation
Supported
Voice cloning
Not supported
Controllability

Designed for long-form audio (audiobooks, podcasts) with full SSML.

Boson AI

Higgs TTS 3

1,500credits
/ 1M chars

Boson's open-weights model, served hosted: 100+ languages detected automatically, zero-shot cloning from a single reference clip, and inline tags for emotion, style and sound effects.

Voices
6 voices
Languages
+311 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Inline tags for emotion, style (whisper, shout, sing), sound effects, pace and pauses. No numeric speed or pitch, no SSML.

CAMB.AI

MARS Pro

10,256credits
/ 1M chars

Massively multilingual TTS (140+ languages) with cross-lingual voice transfer and reference-audio emotion for global dubbing.

Voices
793 voices
Languages
+210 regions supported
Custom pronunciation
Not supported
Voice cloning
Supported
Controllability

Reference-audio emotion, speed, age and gender, plus cross-lingual voice transfer.

Cartesia

Sonic 3.5

3,920credits
/ 1M chars

Cartesia's May-2026 flagship: #1 on the Artificial Analysis Speech Arena with ~82ms latency, 50+ emotions and IPA dictionaries.

Voices
594 voices
Languages
+311 regions supported
Custom pronunciation
Supported
Voice cloning
Supported
Controllability

Speed and volume, 50+ emotions, <break> and <spell> tags and pronunciation dictionaries.

Cartesia

Sonic 3.6

3,920credits
/ 1M chars

Cartesia's August-2026 stable release, preferred over Sonic 3.5 nearly two to one in blind listening: context-aware pacing without tags, native alphanumerics, and 44 languages on the same voices.

Voices
594 voices
Languages
+311 regions supported
Custom pronunciation
Supported
Voice cloning
Supported
Controllability

Speed and volume, 60 emotion values (beta), <break> and <spell> tags, [laughter], and IPA or sounds-like pronunciation dictionaries.

Cartesia

Sonic Preview

3,920credits
/ 1M chars

Cartesia's rolling beta channel: it always points at the next Sonic build ahead of the stable release, so you hear upcoming changes early. It can change without notice and is meant for testing, not production.

Voices
594 voices
Languages
+311 regions supported
Custom pronunciation
Supported
Voice cloning
Supported
Controllability

Same levers as the current Sonic release (speed, volume, emotion, <break>, <spell>, pronunciation dictionaries), subject to change between builds.

Deepgram

Aura-2

3,000credits
/ 1M chars

Sub-200ms conversational English TTS built for voice agents, with SSML phoneme control and a natural, on-brand delivery.

Voices
50 voices
Languages
4 regions supported
Custom pronunciation
Supported
Voice cloning
Not supported
Controllability

Speed, pitch and volume via SSML, emphasis tags and a conversational style.

Deepgram

Flux TTS

4,500credits
/ 1M chars

English TTS for voice agents that holds dialogue context across turns and reports what was already spoken when a caller cuts in. First audio lands in ~80ms.

Voices
36 voices
Languages
2 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Speed 0.5 to 1.5 in 0.05 steps and an expressivity dial (beta), with no SSML or pause tags.

ElevenLabs

Eleven Flash v2.5

5,000credits
/ 1M chars

Lowest-latency ElevenLabs model (~75ms) for real-time agents, with strong Chinese and Cantonese and 300+ voices.

Voices
732 voices
Languages
+311 regions supported
Custom pronunciation
Not supported
Voice cloning
Supported
Controllability

Stability, similarity and speed (limited expression).

ElevenLabs

Eleven Multilingual v2

10,000credits
/ 1M chars

Studio-grade multilingual speech with wide locale coverage and full SSML. Best-in-class Chinese and Cantonese.

Voices
732 voices
Languages
+311 regions supported
Custom pronunciation
Not supported
Voice cloning
Supported
Controllability

Stability, similarity, style exaggeration and speed (0.7 to 1.2).

ElevenLabs

Eleven Turbo v2.5

5,000credits
/ 1M chars

Balanced quality and speed with audio markup tags and instant cloning; sub-250ms latency for interactive use.

Voices
732 voices
Languages
+311 regions supported
Custom pronunciation
Not supported
Voice cloning
Supported
Controllability

Stability, similarity, style and speed.

ElevenLabs

Eleven v3

10,000credits
/ 1M chars

Alpha flagship with inline audio-tag emotion ([whispering], [laughing]) across 70+ languages. The most expressive ElevenLabs model.

Voices
732 voices
Languages
8 regions supported
Custom pronunciation
Supported
Voice cloning
Supported
Controllability

Audio tags ([whispering], [laughing], [excited]) plus stability, similarity and style exaggeration.

Fish Audio

S2 Pro

1,500credits
/ 1M bytes

Expressive successor to S1 with open-ended [bracket] emotion tags, 64+ expressions and ~100ms latency; strong CJK coverage.

Voices
839 voices
Languages
8 regions supported
Custom pronunciation
Not supported
Voice cloning
Supported
Controllability

Open-ended [bracket] emotion tags, 64+ expressions and paralinguistics (laugh, sigh, whisper).

Google Cloud

Gemini 3.1 Flash TTS

3,358credits
/ 1M chars

Gemini-native TTS with natural-language style prompts and 200+ expressive audio tags across single- and multi-speaker output.

Voices
30 voices
Languages
7 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Natural-language style prompts and 200+ expressive audio tags ([whispers], [laughs]).

Google Cloud

WaveNet voices

400credits
/ 1M chars

DeepMind WaveNet voices across 40+ languages with full SSML control. The workhorse for high-volume, multilingual pipelines.

Voices
40 voices
Languages
+19 regions supported
Custom pronunciation
Supported
Voice cloning
Not supported
Controllability

SSML prosody (rate/pitch/volume), emphasis, breaks, say-as and phoneme tags.

Gradium

Gradium TTS

4,778credits
/ 1M chars

Gradium's TTS for voice agents: median first audio under 250ms, instant cloning from a 10-second clip and word-level timestamps for tight sync. English, French, German, Spanish and Portuguese.

Voices
386 voices
Languages
5 regions supported
Custom pronunciation
Not supported
Voice cloning
Supported
Controllability

Speed, randomness (temperature) and voice similarity, <break> tags up to 2 seconds, and a pronunciation dictionary.

Inworld

Realtime TTS 1.5 Max

3,500credits
/ 1M chars

The Max tier of Inworld's 1.5 line, with inline IPA control and expressive non-verbal markers for real-time agents.

Voices
202 voices
Languages
+311 regions supported
Custom pronunciation
Supported
Voice cloning
Supported
Controllability

Inline IPA, non-verbal markers and emphasis (full NL steering on TTS-2).

Inworld

Realtime TTS-2

2,500credits
/ 1M chars

Inworld's research-preview flagship (May 2026), #1 on the Artificial Analysis Speech Arena, with LLM-style natural-language steering.

Voices
202 voices
Languages
+311 regions supported
Custom pronunciation
Supported
Voice cloning
Supported
Controllability

Natural-language steering in [brackets], three stability modes and audio-context conditioning.

Microsoft Azure

Azure HD voices

2,200credits
/ 1M chars

Azure's HD tier: natural, conversational delivery that adapts tone in context, with full SSML and custom lexicon support.

Voices
25 voices
Languages
6 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Natural conversational tone with speed and full SSML.

Microsoft Azure

Azure Neural TTS

1,500credits
/ 1M chars

Azure's neural voices with the broadest language coverage in the market: full SSML, emotional styles and custom lexicons.

Voices
275 voices
Languages
+311 regions supported
Custom pronunciation
Supported
Voice cloning
Not supported
Controllability

Full SSML: prosody, emotional styles (cheerful/sad/angry) and viseme output.

Microsoft Azure

MAI-Voice-2

2,200credits
/ 1M chars

Microsoft's newest MAI flagship TTS (Build 2026 preview) with expressive styles, multi-speaker and long-form via Azure Speech SSML.

Voices
22 voices
Languages
+19 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Expressive styles with style degree, 15+ emotions, multi-speaker and long-form.

MiniMax

Speech 2.8 HD

10,000credits
/ 1M chars

Multilingual HD flagship (Jan 2026) with deep controllability: emotion, voice mixing, pinyin/Jyutping tones and IPA dictionaries.

Voices
332 voices
Languages
+311 regions supported
Custom pronunciation
Supported
Voice cloning
Supported
Controllability

Emotion, speed/pitch/volume, voice mixing, pause and interjection tags, plus IPA dictionaries.

Murf AI

Falcon 2

1,000credits
/ 1M chars

Studio-friendly voices for content teams, with a pronunciation library, IPA overrides and simple speed, pitch and emphasis controls.

Voices
78 voices
Languages
+210 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Speed, pitch, emphasis, pause and pronunciation overrides.

Naver

Clova

7,093credits
/ 1M chars

Korea's market-leading TTS, with premium Korean voices plus English, Japanese and Chinese, avatar video and cloning.

Voices
89 voices
Languages
5 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Emotion (where supported), speed, pitch, volume and pause.

OpenAI

GPT-4o mini TTS

2,000credits
/ 1M chars

Steerable OpenAI voices driven by a natural-language instructions field ('speak slowly and warmly'). Simple, fast and inexpensive.

Voices
13 voices
Languages
8 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Natural-language instructions field ('speak slowly and warmly') plus speed.

Respeecher

Real-Time TTS

3,400credits
/ 1M chars

Studio-grade, ethically-sourced voice cloning used by EA, Sony and Lucasfilm; speech-to-speech that captures a real performance.

Voices
25 voices
Languages
1 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Pitch, emotion and formant driven by a reference performance (speech-to-speech), plus de-noising.

Respeecher Marketplace

Respeecher Marketplace TTS

7,450credits
/ 1M chars

Licensed voices from real actors on Respeecher's Voice Marketplace, each with its own narration styles, and the talent is paid on every use. Rendered per order, not streamed: built for narration and dubbing, not live agents.

Voices
63 voices
Languages
+311 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Per-voice narration style and language. No speed, pitch or SSML.

Rime

Coda

3,000credits
/ 1M chars

A sub-100ms naturalness leader (Coval bench) tuned for effortless, hands-off conversational speech, with deliberately minimal controls.

Voices
223 voices
Languages
+19 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Minimal by design: playback speed and voice selection only.

Rime

Mist v2

3,000credits
/ 1M chars

Low-latency conversational model with 300+ voices and deterministic IVR pronunciation for high-volume agents.

Voices
129 voices
Languages
6 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Speed, deterministic IVR pronunciation and pause control.

Rime

Mist v3

3,000credits
/ 1M chars

Rime's enterprise conversational model with time to first byte under 100ms and an engine rebuilt for concurrent load, on the same voices as Mist v2 across English, Spanish, German and French.

Voices
71 voices
Languages
5 regions supported
Custom pronunciation
Not supported
Voice cloning
Not supported
Controllability

Speed (whole line or per bracketed word), custom pauses in brackets and spell() for letter-by-letter reading. No inline pronunciation.

Don’t see your model?

Ready to publish?

Turn any script into production-quality voice,
in any language, in minutes.

Run your first line