Onepin launches August 11, 9 AM PT.

← Back to blog
Jun 19, 2026

ZONOS2: The First Open-Source MoE TTS Model and the Production Problem It Creates

Zyphra's ZONOS2 is the first open-source mixture-of-experts (MoE) text-to-speech model released under the Apache 2.0 license, with 8 billion parameters and two distinct generation modes: Stable and Expressive. It is a genuine step forward for open-source TTS, and it introduces a routing decision that most production teams are not prepared to make at scale.

The Two-Mode Routing Problem

ZONOS2 ships two modes. Stable mode produces consistent, low-variance output — correct for narration, e-learning, IVR. Expressive mode produces higher emotional variance — right for drama and character dialogue, but harder to validate at scale. Most production content pipelines need both. Routing between them per content type, validating against mode-specific quality criteria, and managing retries differently per mode is a three-layer orchestration problem, not a one-time model selection decision.

How Onepin Addresses It

Onepin connects to ZONOS2 alongside 100+ other TTS models including ElevenLabs, Cartesia, Deepgram Aura-2, and Rime AI. You define which content types route to ZONOS2 Stable, which route to Expressive, and what validation thresholds apply. Onepin runs the jobs, validates, retries, and delivers publish-ready audio. When a better open-source model releases, you add a routing rule — not rebuild your integration. Get started with the Python SDK via pip install onepin, or read the Onepin documentation. onepin.ai

Frequently asked questions

What is ZONOS2?
ZONOS2 is Zyphra's open-source mixture-of-experts text-to-speech model, released under the Apache 2.0 license with 8 billion parameters. The post calls it the first open-source MoE TTS model and notes it ships two distinct generation modes, Stable and Expressive.
What is the difference between ZONOS2 Stable and Expressive modes?
Stable mode produces consistent, low-variance output suited to narration, e-learning, and IVR, while Expressive mode produces higher emotional variance for drama and character dialogue but is harder to validate at scale. Most production pipelines need both.
Why does ZONOS2 create a routing problem for production teams?
Because most content pipelines need both modes, teams must route between them per content type, validate against mode-specific quality criteria, and manage retries differently per mode. The post frames this as a three-layer orchestration problem rather than a one-time model selection decision.
How does Onepin handle ZONOS2 in production?
Onepin connects to ZONOS2 alongside 100+ other TTS models including ElevenLabs, Cartesia, Deepgram Aura-2, and Rime AI. You define which content types route to Stable and which to Expressive and what validation thresholds apply, and Onepin runs the jobs, validates, retries, and delivers publish-ready audio, so adding a better model later is a routing rule rather than a rebuild.