JoinInDocs
Baton
v2 (beta)
Get a key
Baton v2 Guides Beta

Rooms

Per-speaker transcripts and turn events from one mixed recording.

/v2/room takes the mixed audio of a meeting room or call, separates up to 8 voices, transcribes each voice on its own, and reports turn events for the room as a whole.

/v2/room is served on baton-test.joinin.ai. Production serves the /v2 endpoints soon.

Start a room

A room takes the same session.start options as the streams: eot, language, signals, idle_timeout_s.

{
  "type": "session.start",
  "api_key": "baton_...",
  "eot": {"profile": "cutoff_5pct"},
  "language": {"allowed": ["en", "es", "fr"]},
  "signals": ["vad", "eot", "bargein"]
}

Then send the mixed audio as binary frames, exactly as on the stream endpoints.

Modes

mode is optional. It only sets the default operating point, used when session.start has no eot; an eot always wins. Valid values:

mode Default operating point Mean latency in a room (English)
transactional_call (used when mode is omitted) cutoff_10pct 439 ms
working_session cutoff_5pct 726 ms
one_on_one cutoff_5pct 726 ms
brainstorm cutoff_2pct 1134 ms
formal cutoff_2pct 1134 ms

Any other value is refused with 1008 unknown mode '<mode>'.

Transcripts

Each voice is named spk_0 to spk_7 in order of first appearance, and keeps its name for the session.

{"type": "asr.partial", "seq": 2, "t": 3.2,
 "speaker": "spk_1", "text": "I", "language": "en"}
{"type": "asr.final", "seq": 4, "t": 3.52,
 "speaker": "spk_1", "text": "I never", "t0": 3.2,
 "t1": 3.52, "language": "en"}
  • asr.partial carries a voice's newest word or words, which may still change. Each partial replaces the previous one for that voice.
  • asr.final carries words that will not change. Append them to that voice's transcript.
  • A word becomes final as soon as the speaker says another word, after about a second without change, or when you send finalize.
  • t0 and t1 are when the first and last of the final words were first heard.

Cutting segments at turn boundaries

Turn events in a room use speaker: "room". When your own logic decides a speaker has finished, send finalize:

{"type": "finalize"}

Baton sends each voice's pending word as asr.final, then asr.flushed. Cut your segments on asr.flushed. Otherwise a speaker's last word can arrive after you have cut, splitting "Saint" from "Paul".

Ending

Send session.end. Baton sends the remaining finals, a language.summary (seconds of speech per voice per language), session.summary (audio_s, turns), then closes with 1000.

Barge-in and idle sessions

Both work as on the streams. Send agent.speech start and end while your agent talks to get bargein.onset and bargein.class when anyone in the room speaks over your agent. A room silent for idle_timeout_s (default 600) gets a warning, then closes with 1000 idle_timeout; its final transcripts are sent first. See Barge-in.

messages is not used in rooms.

© 2026 JoinIn AI, Inc. All rights reserved.joinin.ai