JoinInDocs
Baton
v2 (beta)
Get a key
Baton v2 Streaming (WebSocket) Beta

Room

One mixed recording of several people. Turn events plus per-speaker transcripts.

WSS/v2/room

Send the room's mixed audio as one source. Baton separates the voices (up to 8), transcribes each voice on its own, identifies each voice's language, and reports turn events for the room as a whole (speaker: "room").

Transcripts. Each voice gets asr.partial for its newest, still-changing word and asr.final for words that will not change. A word becomes final as soon as the speaker says another word, after about a second without change, or when you send finalize.

Turn boundaries. When your own turn logic decides a speaker has finished, send finalize. Every voice's pending word is emitted as asr.final, followed by asr.flushed. Cut your segments on asr.flushed, not on turn.end.

Languages. Baton identifies each voice's language continuously (asr.language). When a voice is clearly speaking another language from your allowed list, its transcription switches to that language (language.switched). Without an allowed list Baton only warns (warning with code: language_mismatch).

Same session options as the streams: signals (add bargein for barge-in events when anyone in the room speaks over your agent), agent.speech, idle_timeout_s, eot. messages is not used in rooms.

When rooms are at capacity the start closes with 1013; retry with backoff.

Same as /v2/streamIdentical messages; room events carry speaker: "room".
Show Session, Audio, Speech and turns, Barge-in

Session

Open, configure and close the session.

session.startyou sendThe first frame of a room. The same options as a stream's session.start.

Rejections close with 1008 (invalid api key, unknown mode '<mode>', a language or end-of-turn error, or unsupported language 'xx') or 1013 when rooms are at capacity or temporarily unavailable (retry with backoff).

FieldTypeDescription
type required"session.start"
api_key requiredstring
eotone of: By profile, By latency budget

Choose an end-of-turn operating point by name, or by the mean latency you can afford (the closest published point wins; ties go to the faster one). Give at most one; omit eot for cutoff_10pct.

Operating point Mean latency, /v2/stream Mean latency, /v2/room False cut-offs
cutoff_10pct 350 ms 439 ms about 10 %
cutoff_5pct 577 ms 726 ms about 5 %
cutoff_2pct 875 ms 1134 ms about 2 %

English figures; every language's points are at GET /v2/eot/profiles. Examples: {"profile": "cutoff_5pct"}, or {"latency_budget_ms": 600} for the point whose mean latency is closest to 600 ms.

languageone of: Pinned, Allowed set, Automatic

Default {"pin": "en"}.

signalsarray of "vad" | "eot" | "bargein"

VAD and end of turn are always on. Add bargein to receive barge-in events.

idle_timeout_snumber

Close after this many seconds without speech. 0 or less disables it. Default 600.

mode"transactional_call" | "working_session" | "one_on_one" | "brainstorm" | "formal"

Default "transactional_call".

inference_intervalnumber

Default 0.1.

An English meeting that may switch to Spanish or French speakers.

{
  "type": "session.start",
  "api_key": "baton_0123456789abcdef0123456789abcdef0123456789abcdef",
  "eot": {"profile": "cutoff_5pct"},
  "language": {"allowed": ["en", "es", "fr"]},
  "signals": ["vad", "eot", "bargein"],
  "idle_timeout_s": 900
}

session.startedyou receiveThe room is open and the transcriber is ready.

FieldTypeDescription
type required"session.started"
schema_version required2
session_id requiredstring
eot requiredobject

The operating point this session uses.

modelsobject

Present when bargein was requested.

{
  "type": "session.started",
  "schema_version": 2,
  "session_id": "8b1e4f2a9c3d4e7f8a1b2c3d4e5f6a7b",
  "eot": {
    "profile": "cutoff_5pct",
    "mean_latency_ms": 726,
    "cutoff_rate": 0.0482
  }
}

warning (idle_timeout)you receiveNo speech for a while. The session will close soon.

Sent before an idle close (idle_timeout_s, default 600 s). Speech resets the timer. The close is code 1000 with reason idle_timeout.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"warning"
code required"idle_timeout"
close_in_s requirednumber

Seconds until the idle close.

{
  "type": "warning",
  "seq": 40,
  "schema_version": 2,
  "code": "idle_timeout",
  "close_in_s": 30.0
}

session.endyou sendEnd the session. Baton replies with session.summary and closes with 1000.

FieldTypeDescription
type required"session.end"
{"type": "session.end"}

session.summaryyou receiveTotals for the session. Sent in reply to session.end.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"session.summary"
audio_s requirednumber

Seconds of audio received.

turns requiredinteger

Turn ends

{
  "type": "session.summary",
  "seq": 41,
  "schema_version": 2,
  "audio_s": 312.4,
  "turns": 18
}

Audio

16 kHz mono PCM16 in binary frames.

audioyou sendOne binary frame of audio.

A binary WebSocket frame: one leading byte (send 0; it identifies the source and is reserved for multi-source sessions), then signed 16-bit little-endian PCM, mono, 16 000 Hz. Any whole number of samples per frame; 80–100 ms (1280–1600 samples) is typical. Resample on the client: other rates are not accepted.

Binary frame. [u8 source = 0][PCM16LE mono 16 kHz samples...]

Speech and turns

When people speak, pause and finish.

vad.startyou receiveSomeone in the room started speaking.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"vad.start"
tnumber

Audio time in seconds since the session's first sample.

speaker"room"

vad.endyou receiveThe room went quiet.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"vad.end"
tnumber

Audio time in seconds since the session's first sample.

speaker"room"
durationnumber

turn.endyou receiveThe room's current turn has ended.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"turn.end"
tnumber

Audio time in seconds since the session's first sample.

speaker"room"
p_eotnumber
latency_msinteger
by"score" | "timeout"
profilestring

turn.resumedyou receiveSpeech resumed within 2 s of turn.end.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"turn.resumed"
tnumber

Audio time in seconds since the session's first sample.

speaker"room"
after_msinteger

Barge-in

Speech over your agent: a real interruption, or a backchannel.

agent.speechyou sendTell Baton when your agent starts and stops talking.

Needed for barge-in. While the agent is speaking, a user speech onset becomes a bargein.onset and is then classified. Sending phase: end drops an onset that has not been classified yet. Timestamped on arrival against the audio received so far. Ignored unless the session subscribed to bargein.

FieldTypeDescription
type required"agent.speech"
phase required"start" | "end"
utterance_idstring

Your id for the agent utterance. Echoed on barge-in events.

{
  "type": "agent.speech",
  "phase": "start",
  "utterance_id": "sp_1"
}
{
  "type": "agent.speech",
  "phase": "end",
  "utterance_id": "sp_1"
}

bargein.onsetyou receiveThe user started speaking while your agent was talking.

Sent at the speech onset. Classification follows in bargein.class.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"bargein.onset"
t requirednumber

Audio time in seconds since the session's first sample.

agent_speaking requiredtrue
utterance_idstring
{
  "type": "bargein.onset",
  "seq": 7,
  "schema_version": 2,
  "t": 12.1,
  "agent_speaking": true,
  "utterance_id": "sp_1"
}

bargein.classyou receiveWhether the interruption claims the floor or is a backchannel.

claim means the user wants to talk: stop or yield. backchannel ("mm-hm", "right") means keep talking. Decided as soon as Baton is confident, typically a few hundred milliseconds into the interruption. p_claim is null when the user stopped too soon to judge; that case is reported as a backchannel.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"bargein.class"
t requirednumber

Audio time in seconds since the session's first sample.

class required"claim" | "backchannel"
p_claim requirednumber or null
decided_after_ms requiredinteger

Milliseconds from the onset to this decision.

utterance_idstring
{
  "type": "bargein.class",
  "seq": 8,
  "schema_version": 2,
  "t": 12.4,
  "class": "claim",
  "p_claim": 0.9,
  "decided_after_ms": 300,
  "utterance_id": "sp_1"
}
Added in roomsWhat a room adds on top of the stream.

Transcripts

Per-speaker words, and clean cuts at turn boundaries.

finalizeyou sendFinalise every voice's pending word now (a turn boundary you decided).

Baton emits each voice's pending newest word as asr.final, then asr.flushed. Without it, a speaker's last word waits up to 1 s and can land after you have already cut the segment ("...in Saint" / "Paul").

FieldTypeDescription
type required"finalize"
{"type": "finalize"}

asr.partialyou receiveA voice's newest word or words, still subject to change.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"asr.partial"
t requirednumber

Audio time in seconds since the session's first sample.

speaker requiredstring

A voice in the room, numbered in order of first appearance. Stable for the session.

text requiredstring

The voice's newest word or words not yet final. Replaces the previous partial for this voice.

language requiredstring

The language this voice is decoded in.

{
  "type": "asr.partial",
  "seq": 2,
  "schema_version": 2,
  "t": 3.2,
  "speaker": "spk_1",
  "text": "I",
  "language": "en"
}

asr.finalyou receiveWords from one voice that will not change.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"asr.final"
t requirednumber

Audio time in seconds since the session's first sample.

speaker requiredstring

A voice in the room, numbered in order of first appearance. Stable for the session.

text requiredstring

Final words. Append to the voice's transcript.

t0 requirednumber

When the first of these words was first heard (seconds).

t1 requirednumber

When the last of these words was first heard (seconds).

language requiredstring
{
  "type": "asr.final",
  "seq": 4,
  "schema_version": 2,
  "t": 3.52,
  "speaker": "spk_1",
  "text": "I never",
  "t0": 3.2,
  "t1": 3.52,
  "language": "en"
}

asr.flushedyou receiveEvery word pending at your finalize has been sent. Safe to cut segments.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"asr.flushed"
{"type": "asr.flushed", "seq": 6, "schema_version": 2}

Languages

Each voice's language, and switching within your allowed set.

asr.languageyou receiveLanguage identified on one voice's last 3 s of solo speech.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"asr.language"
t requirednumber

Audio time in seconds since the session's first sample.

speaker requiredstring

A voice in the room, numbered in order of first appearance. Stable for the session.

language requiredstring
confidence requirednumber
{
  "type": "asr.language",
  "seq": 30,
  "schema_version": 2,
  "t": 15.36,
  "speaker": "spk_0",
  "language": "en",
  "confidence": 0.9981
}

warning (language_mismatch)you receiveA voice sounds like a different language and the room has no allowed list.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"warning"
code"language_mismatch"
tnumber

Audio time in seconds since the session's first sample.

speakerstring

A voice in the room, numbered in order of first appearance. Stable for the session.

languagestring

language.switchedyou receiveA voice is now decoded in another language from your allowed list.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"language.switched"
tnumber

Audio time in seconds since the session's first sample.

speakerstring

A voice in the room, numbered in order of first appearance. Stable for the session.

languagestring
{
  "type": "language.switched",
  "seq": 52,
  "schema_version": 2,
  "t": 41.6,
  "speaker": "spk_2",
  "language": "es"
}

language.switch_recommendedyou receiveA voice needs a language this room cannot switch to by itself.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"language.switch_recommended"
tnumber

Audio time in seconds since the session's first sample.

speakerstring

A voice in the room, numbered in order of first appearance. Stable for the session.

languagestring

language.summaryyou receiveSeconds of solo speech per voice per language. Sent at the end.

FieldTypeDescription
seq requiredinteger

Event sequence number. Starts at 1 and increases by one per event.

schema_version required2
type required"language.summary"
tnumber

Audio time in seconds since the session's first sample.

speakersobject

{voice: {language: seconds of solo speech}}

{
  "type": "language.summary",
  "seq": 900,
  "schema_version": 2,
  "t": 1802.24,
  "speakers": {
    "spk_0": {"en": 512.0},
    "spk_1": {"en": 301.0, "es": 12.0}
  }
}
© 2026 JoinIn AI, Inc. All rights reserved.joinin.ai