Skip to navigation

Fix pronunciations

One wrong name ruins the take — catch it in the run, and record how it should sound.

TTS models guess at names, brands, and jargon — and one wrong name ruins the take. Onepin gives you two separate things, and it’s worth being clear about which does what:

  1. Inside a run — a chain of three nodes that catches a mispronunciation in the generated audio and fixes it. This is what actually changes what you hear.
  2. In your workspace — a dictionary that records how your team has agreed a term should sound. This is a record, not yet an input to a run.

Catch and fix it in the run

Three nodes work together, all en-us / en-gb only and all in beta:

Generator ──▶ Phoneme Injector ──▶ Pronunciation ──pass──▶ Export
│
fail
▼
Pronunciation Corrector
│
└──▶ back to Pronunciation (re-grade)
  • The Phoneme Injector reads the script, picks out the names and terms worth checking, and looks up how each should sound in Onepin’s built-in pronunciation reference.
  • The Pronunciation validator listens to the generated audio and scores it phoneme by phoneme against those references. Its default bar is 99 — near zero defects.
  • The Pronunciation Corrector re-synthesizes just the words that were wrong, inside the existing take, keeping the voice and the surrounding audio untouched. That beats sending the line back to the Generator for a whole new read.

Set the bar on the check like any other validator:

for node in definition["graph"]["nodes"]:
if node["type"] == "validator_pronunciation":
node["config"] = {"threshold": 99.0, "max_retries": 3}

Placement is strict — the Injector must sit immediately before the check, the check goes last among a branch’s validators, and the Corrector hangs off the check’s fail (or pass) port. See General rules for what the graph validator enforces.

On a non-English branch, or when you’d rather not add the chain, the Accuracy validator is the coarser safety net: it hears a bad mispronunciation as a word miss and regenerates the line.

Record it in the workspace dictionary

A dictionary entry is your workspace’s agreed spelling-to-sound record for a term — one per language, shared by everyone. Needs the dictionary:read / dictionary:write scopes.

Add an entry

Two methods:

  • spelled — a phonetic respelling (“Podonos” → “poh-DOH-nohs”)
  • recorded — reference audio of the word, from an upload
client.dictionary.create_dictionary_entry(
word="Podonos",
method="spelled",
pronunciation="poh-DOH-nohs",
language="en-us",
)

Not sure how to respell it? Ask for a suggestion first:

s = client.dictionary.suggest_pronunciation(word="Podonos", language="en-us").data

Entries are per-language — a name said one way in English and another in Korean gets one entry per language.

Review what’s in the dictionary

for e in client.dictionary.list_dictionary_entries(language="en-us").data:
print(e.word, "→", e.pronunciation)
client.dictionary.search_dictionary_entries(search="podo")
client.dictionary.list_dictionary_languages() # languages you can add entries for

Dictionary management is SDK/API only for now — the CLI doesn’t cover it yet.

The dictionary does not feed the run yet. The pronunciation guidance the Phoneme Injector uses comes from Onepin’s built-in reference, not from entries you add here. Your entries are a shared record for your team; they are not injected into synthesis or into the checks.