/v2/room takes the mixed audio of a meeting room or call, separates up to 8 voices, transcribes each voice on its own, and reports turn events for the room as a whole.
/v2/room is served on baton-test.joinin.ai. Production serves the /v2 endpoints soon.
Start a room
A room takes the same session.start options as the streams: eot, language, signals, idle_timeout_s.
{
"type": "session.start",
"api_key": "baton_...",
"eot": {"profile": "cutoff_5pct"},
"language": {"allowed": ["en", "es", "fr"]},
"signals": ["vad", "eot", "bargein"]
}
Then send the mixed audio as binary frames, exactly as on the stream endpoints.
Modes
mode is optional. It only sets the default operating point, used when session.start has no eot; an eot always wins. Valid values:
mode |
Default operating point | Mean latency in a room (English) |
|---|---|---|
transactional_call (used when mode is omitted) |
cutoff_10pct |
439 ms |
working_session |
cutoff_5pct |
726 ms |
one_on_one |
cutoff_5pct |
726 ms |
brainstorm |
cutoff_2pct |
1134 ms |
formal |
cutoff_2pct |
1134 ms |
Any other value is refused with 1008 unknown mode '<mode>'.
Transcripts
Each voice is named spk_0 to spk_7 in order of first appearance, and keeps its name for the session.
{"type": "asr.partial", "seq": 2, "t": 3.2,
"speaker": "spk_1", "text": "I", "language": "en"}
{"type": "asr.final", "seq": 4, "t": 3.52,
"speaker": "spk_1", "text": "I never", "t0": 3.2,
"t1": 3.52, "language": "en"}
asr.partialcarries a voice's newest word or words, which may still change. Each partial replaces the previous one for that voice.asr.finalcarries words that will not change. Append them to that voice's transcript.- A word becomes final as soon as the speaker says another word, after about a second without change, or when you send
finalize. t0andt1are when the first and last of the final words were first heard.
Cutting segments at turn boundaries
Turn events in a room use speaker: "room". When your own logic decides a speaker has finished, send finalize:
{"type": "finalize"}
Baton sends each voice's pending word as asr.final, then asr.flushed. Cut your segments on asr.flushed. Otherwise a speaker's last word can arrive after you have cut, splitting "Saint" from "Paul".
Ending
Send session.end. Baton sends the remaining finals, a language.summary (seconds of speech per voice per language), session.summary (audio_s, turns), then closes with 1000.
Barge-in and idle sessions
Both work as on the streams. Send agent.speech start and end while your agent talks to get bargein.onset and bargein.class when anyone in the room speaks over your agent. A room silent for idle_timeout_s (default 600) gets a warning, then closes with 1000 idle_timeout; its final transcripts are sent first. See Barge-in.
messages is not used in rooms.