Get Voice Facets
Filter-bar options (chips) for the voice browser, one list per dimension.
Returns providers, models, languages (data-driven) plus genders,
ages, categories, accents (fixed enums) as VoiceFacetItem[] so the FE
builds the whole filter bar — with per-chip count badges — in a single request
instead of hardcoding option lists (mirrors GET /dictionary/languages). Each
item is {value, label, count}: value is passed straight back to
GET /voices; label is the display name for providers/models and null
elsewhere (the FE owns language + enum labels); count is the number of
matching voices. A language surfaces as one chip keyed by its canonical
allowlist locale: every declared locale is folded onto the locale the
GET /voices filter would match it against (bare ko and regioned ko-kr
both count under ko-kr), so variants never split into duplicate chips for
the identical filter. Model counts include only voices that explicitly declare
the model. Language counts include voices that declare the exact regional locale
plus voices that declare its bare family (en contributes to every supported
en-* locale). A voice with no declared supported_models is “general use” —
GET /voices matches it against every model filter but no models chip counts
it, so a model chip’s count can be lower than the GET /voices?model= result.
Languages have no such gap: no-locale voices are excluded from both the language
chips and GET /voices?language=, so language chip counts match the row counts.
Accepts the SAME filters as GET /voices (tab scope source/favorites_only,
plus provider/model/language/gender/age/category/accent/search).
count is context-aware (faceted search): each dimension’s counts apply every
OTHER active filter but exclude that dimension’s own selection — e.g. with
provider=elevenlabs the language counts are scoped to ElevenLabs, while the
provider chips still show every provider so the caller can switch.
search follows the SAME two modes as GET /voices (see _semantic_search_active).
In SEMANTIC mode (flag on, embedder available, platform in scope) the chips are
counted over the population the semantic list can return — every active filter, plus
“carries a description embedding OR is in the current ranked set” (the ANN only ranks
embedded voices; the keyword arm contributes the rest) — so the chips describe the
voices the list actually shows and never collapse to “No matches” on a query that has
no literal name/tag hit. Per-dimension self-exclusion applies in full, exactly as in
lexical mode. Counts are clamped to VOICE_SEARCH_MAX_RANKED, because the semantic
list’s pagination.total is that same capped ranked-set size — so a chip’s count is
the number of rows GET /voices returns once that value is selected. Two bounded
exceptions: a voice reachable only through the keyword arm is counted only under the
value already selected (the ranked set is computed under the current filters), and
ANN recall can return fewer rows than the chip promises. In the LEXICAL
fallback (flag off / no embedder / non-platform source / embed fault) search is the
name/descriptor/tags ILIKE and counts are exact (no cap), exactly as before.
Chips are drawn from the same population GET /voices returns, so the
official-locale restriction applies here too and no chip can open an empty page.
For the model dimension the restriction is evaluated per capability row rather
than per voice — a voice whose only official locale sits on a sibling model does
not count toward this model’s chip, because GET /voices?model= would not return
it either.
Count-0 policy: data-driven dimensions omit count-0 values (only present ones,
each a valid GET /voices filter — providers/models restricted to the enabled
catalog, languages to the supported-locale allowlist, so a chip never 422s).
Enum dimensions always return the full enum in natural order, count-0 included,
for the FE to grey out.
Authentication
Onepin live API key (op_live_...). Test and public keys are reserved in Phase 1.
Headers
Query parameters
Tab scope — repeat for OR, same values as GET /voices (e.g. platform, workspace)
Repeat for OR, e.g. ?provider=elevenlabs&provider=rime
Repeat for OR. Filters platform voices by TTS model, e.g. ?model=arcana&model=sonic-2
Repeat for OR, e.g. ?language=en-us&language=ko-kr
Response
Filter options for the voice browser, one VoiceFacetItem[] per chip.
Two families of dimension:
- Data-driven —
providers,models,languages: only values with a scoped voice count are returned (count is always ≥ 1; count-0 values are omitted). A language value may be derived from a bare family tag on a scoped voice. Every value is guaranteed to be a validGET /voicesfilter (provider/model restricted to the enabled catalog, language to the supported-locale allowlist), so selecting one never yields a 422 or empty page. Sorted count DESC, then value ASC. - Enum —
genders,ages,categories,accents: the FULL fixed enum is always returned in natural enum order, including count-0 values (the FE greys those out).labelisNone(the FE owns enum labels).
count is context-aware (faceted search): each dimension's counts apply
every OTHER active filter but exclude that dimension's own selection, so a
chip's number reflects "results if I also pick this" without the dimension
suppressing its own alternatives.

