language in session.start is always an object. Codes are BCP-47 style: en, es, ja-JP.
| Policy | Example | Behaviour |
|---|---|---|
| Pin | {"pin": "en"} (default) |
One language for the session. {"pin": "auto"} detects. |
| Allowed set | {"allowed": ["en", "es", "fr"]} |
The first is the starting language. Rooms switch individual voices within the set. |
| Automatic | {"mode": "auto"} |
Detect the language. |
End-of-turn operating points are published for 14 languages: ar, de, en, es, fr, hi, id, it, ja, ko, nl, pt, tr, zh. Other languages use the English points.
Rooms: per-voice language
Baton identifies the language of each voice continuously and reports it:
{
"type": "asr.language",
"seq": 30,
"t": 15.36,
"speaker": "spk_0",
"language": "en",
"confidence": 0.9981
}
When a voice is clearly speaking another language from your allowed set, Baton switches that voice's transcription to it:
{
"type": "language.switched",
"seq": 52,
"t": 41.6,
"speaker": "spk_2",
"language": "es"
}
Without an allowed set, Baton does not switch. It sends a warning instead:
{
"type": "warning",
"code": "language_mismatch",
"t": 41.6,
"speaker": "spk_2",
"language": "es"
}
language.switch_recommended means a voice needs a language the room cannot switch to by itself.
For English-led meetings that may include other speakers, use an allowed set led by English. Transcription is most accurate when the room starts in the language most people speak.