WSS/v2/room
Send the room's mixed audio as one source. Baton separates the voices (up to 8),
transcribes each voice on its own, identifies each voice's language, and reports turn
events for the room as a whole (speaker: "room").
Transcripts. Each voice gets asr.partial for its newest, still-changing word and
asr.final for words that will not change. A word becomes final as soon as the
speaker says another word, after about a second without change, or when you send
finalize.
Turn boundaries. When your own turn logic decides a speaker has finished, send
finalize. Every voice's pending word is emitted as asr.final, followed by
asr.flushed. Cut your segments on asr.flushed, not on turn.end.
Languages. Baton identifies each voice's language continuously (asr.language).
When a voice is clearly speaking another language from your allowed list, its
transcription switches to that language (language.switched). Without an allowed
list Baton only warns (warning with code: language_mismatch).
Same session options as the streams: signals (add bargein for barge-in
events when anyone in the room speaks over your agent), agent.speech,
idle_timeout_s, eot.
messages is not used in rooms.
When rooms are at capacity the start closes with 1013; retry with backoff.
speaker: "room".Show Session, Audio, Speech and turns, Barge-in
Session
Open, configure and close the session.
session.startyou sendThe first frame of a room. The same options as a stream's session.start.
Rejections close with 1008 (invalid api key, unknown mode '<mode>', a
language or end-of-turn error, or unsupported language 'xx') or 1013 when
rooms are at capacity or temporarily unavailable (retry with backoff).
| Field | Type | Description | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
type required | "session.start" | |||||||||||||||||
api_key required | string | |||||||||||||||||
eot | one of: By profile, By latency budget | Choose an end-of-turn operating point by name, or by the mean latency you can
afford (the closest published point wins; ties go to the faster one). Give at most
one; omit
English figures; every language's points are at | ||||||||||||||||
language | one of: Pinned, Allowed set, Automatic | Default | ||||||||||||||||
signals | array of "vad" | "eot" | "bargein" | VAD and end of turn are always on. Add | ||||||||||||||||
idle_timeout_s | number | Close after this many seconds without speech. 0 or less disables it. Default | ||||||||||||||||
mode | "transactional_call" | "working_session" | "one_on_one" | "brainstorm" | "formal" | Default | ||||||||||||||||
inference_interval | number | Default |
An English meeting that may switch to Spanish or French speakers.
{
"type": "session.start",
"api_key": "baton_0123456789abcdef0123456789abcdef0123456789abcdef",
"eot": {"profile": "cutoff_5pct"},
"language": {"allowed": ["en", "es", "fr"]},
"signals": ["vad", "eot", "bargein"],
"idle_timeout_s": 900
}
session.startedyou receiveThe room is open and the transcriber is ready.
| Field | Type | Description |
|---|---|---|
type required | "session.started" | |
schema_version required | 2 | |
session_id required | string | |
eot required | object | The operating point this session uses. |
models | object | Present when |
{
"type": "session.started",
"schema_version": 2,
"session_id": "8b1e4f2a9c3d4e7f8a1b2c3d4e5f6a7b",
"eot": {
"profile": "cutoff_5pct",
"mean_latency_ms": 726,
"cutoff_rate": 0.0482
}
}
warning (idle_timeout)you receiveNo speech for a while. The session will close soon.
Sent before an idle close (idle_timeout_s, default 600 s). Speech resets the
timer. The close is code 1000 with reason idle_timeout.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "warning" | |
code required | "idle_timeout" | |
close_in_s required | number | Seconds until the idle close. |
{
"type": "warning",
"seq": 40,
"schema_version": 2,
"code": "idle_timeout",
"close_in_s": 30.0
}
session.endyou sendEnd the session. Baton replies with session.summary and closes with 1000.
| Field | Type | Description |
|---|---|---|
type required | "session.end" |
{"type": "session.end"}
session.summaryyou receiveTotals for the session. Sent in reply to session.end.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "session.summary" | |
audio_s required | number | Seconds of audio received. |
turns required | integer | Turn ends |
{
"type": "session.summary",
"seq": 41,
"schema_version": 2,
"audio_s": 312.4,
"turns": 18
}
Audio
16 kHz mono PCM16 in binary frames.
audioyou sendOne binary frame of audio.
A binary WebSocket frame: one leading byte (send 0; it identifies the source
and is reserved for multi-source sessions), then signed 16-bit little-endian PCM,
mono, 16 000 Hz. Any whole number of samples per frame; 80–100 ms (1280–1600
samples) is typical. Resample on the client: other rates are not accepted.
Binary frame. [u8 source = 0][PCM16LE mono 16 kHz samples...]
Speech and turns
When people speak, pause and finish.
vad.startyou receiveSomeone in the room started speaking.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "vad.start" | |
t | number | Audio time in seconds since the session's first sample. |
speaker | "room" |
vad.endyou receiveThe room went quiet.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "vad.end" | |
t | number | Audio time in seconds since the session's first sample. |
speaker | "room" | |
duration | number |
turn.endyou receiveThe room's current turn has ended.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "turn.end" | |
t | number | Audio time in seconds since the session's first sample. |
speaker | "room" | |
p_eot | number | |
latency_ms | integer | |
by | "score" | "timeout" | |
profile | string |
turn.resumedyou receiveSpeech resumed within 2 s of turn.end.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "turn.resumed" | |
t | number | Audio time in seconds since the session's first sample. |
speaker | "room" | |
after_ms | integer |
Barge-in
Speech over your agent: a real interruption, or a backchannel.
agent.speechyou sendTell Baton when your agent starts and stops talking.
Needed for barge-in. While the agent is speaking, a user speech onset becomes a
bargein.onset and is then classified. Sending phase: end drops an onset that has
not been classified yet. Timestamped on arrival against the audio received so far.
Ignored unless the session subscribed to bargein.
| Field | Type | Description |
|---|---|---|
type required | "agent.speech" | |
phase required | "start" | "end" | |
utterance_id | string | Your id for the agent utterance. Echoed on barge-in events. |
{
"type": "agent.speech",
"phase": "start",
"utterance_id": "sp_1"
}
{
"type": "agent.speech",
"phase": "end",
"utterance_id": "sp_1"
}
bargein.onsetyou receiveThe user started speaking while your agent was talking.
Sent at the speech onset. Classification follows in bargein.class.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "bargein.onset" | |
t required | number | Audio time in seconds since the session's first sample. |
agent_speaking required | true | |
utterance_id | string |
{
"type": "bargein.onset",
"seq": 7,
"schema_version": 2,
"t": 12.1,
"agent_speaking": true,
"utterance_id": "sp_1"
}
bargein.classyou receiveWhether the interruption claims the floor or is a backchannel.
claim means the user wants to talk: stop or yield. backchannel ("mm-hm",
"right") means keep talking. Decided as soon as Baton is confident, typically a few
hundred milliseconds into the interruption. p_claim is null when the user
stopped too soon to judge; that case is reported as a backchannel.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "bargein.class" | |
t required | number | Audio time in seconds since the session's first sample. |
class required | "claim" | "backchannel" | |
p_claim required | number or null | |
decided_after_ms required | integer | Milliseconds from the onset to this decision. |
utterance_id | string |
{
"type": "bargein.class",
"seq": 8,
"schema_version": 2,
"t": 12.4,
"class": "claim",
"p_claim": 0.9,
"decided_after_ms": 300,
"utterance_id": "sp_1"
}
Transcripts
Per-speaker words, and clean cuts at turn boundaries.
finalizeyou sendFinalise every voice's pending word now (a turn boundary you decided).
Baton emits each voice's pending newest word as asr.final, then asr.flushed.
Without it, a speaker's last word waits up to 1 s and can land after you have
already cut the segment ("...in Saint" / "Paul").
| Field | Type | Description |
|---|---|---|
type required | "finalize" |
{"type": "finalize"}
asr.partialyou receiveA voice's newest word or words, still subject to change.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "asr.partial" | |
t required | number | Audio time in seconds since the session's first sample. |
speaker required | string | A voice in the room, numbered in order of first appearance. Stable for the session. |
text required | string | The voice's newest word or words not yet final. Replaces the previous partial for this voice. |
language required | string | The language this voice is decoded in. |
{
"type": "asr.partial",
"seq": 2,
"schema_version": 2,
"t": 3.2,
"speaker": "spk_1",
"text": "I",
"language": "en"
}
asr.finalyou receiveWords from one voice that will not change.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "asr.final" | |
t required | number | Audio time in seconds since the session's first sample. |
speaker required | string | A voice in the room, numbered in order of first appearance. Stable for the session. |
text required | string | Final words. Append to the voice's transcript. |
t0 required | number | When the first of these words was first heard (seconds). |
t1 required | number | When the last of these words was first heard (seconds). |
language required | string |
{
"type": "asr.final",
"seq": 4,
"schema_version": 2,
"t": 3.52,
"speaker": "spk_1",
"text": "I never",
"t0": 3.2,
"t1": 3.52,
"language": "en"
}
asr.flushedyou receiveEvery word pending at your finalize has been sent. Safe to cut segments.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "asr.flushed" |
{"type": "asr.flushed", "seq": 6, "schema_version": 2}
Languages
Each voice's language, and switching within your allowed set.
asr.languageyou receiveLanguage identified on one voice's last 3 s of solo speech.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "asr.language" | |
t required | number | Audio time in seconds since the session's first sample. |
speaker required | string | A voice in the room, numbered in order of first appearance. Stable for the session. |
language required | string | |
confidence required | number |
{
"type": "asr.language",
"seq": 30,
"schema_version": 2,
"t": 15.36,
"speaker": "spk_0",
"language": "en",
"confidence": 0.9981
}
warning (language_mismatch)you receiveA voice sounds like a different language and the room has no allowed list.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "warning" | |
code | "language_mismatch" | |
t | number | Audio time in seconds since the session's first sample. |
speaker | string | A voice in the room, numbered in order of first appearance. Stable for the session. |
language | string |
language.switchedyou receiveA voice is now decoded in another language from your allowed list.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "language.switched" | |
t | number | Audio time in seconds since the session's first sample. |
speaker | string | A voice in the room, numbered in order of first appearance. Stable for the session. |
language | string |
{
"type": "language.switched",
"seq": 52,
"schema_version": 2,
"t": 41.6,
"speaker": "spk_2",
"language": "es"
}
language.switch_recommendedyou receiveA voice needs a language this room cannot switch to by itself.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "language.switch_recommended" | |
t | number | Audio time in seconds since the session's first sample. |
speaker | string | A voice in the room, numbered in order of first appearance. Stable for the session. |
language | string |
language.summaryyou receiveSeconds of solo speech per voice per language. Sent at the end.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "language.summary" | |
t | number | Audio time in seconds since the session's first sample. |
speakers | object | {voice: {language: seconds of solo speech}} |
{
"type": "language.summary",
"seq": 900,
"schema_version": 2,
"t": 1802.24,
"speakers": {
"spk_0": {"en": 512.0},
"spk_1": {"en": 301.0, "es": 12.0}
}
}