WSS/v2/stream
One speaker, typically a voice agent's caller. It can use the conversation so far
(messages) to judge whether an utterance is complete. No transcript is returned;
use /v2/room when you need text.
English operating points on this tier:
| profile | mean latency | false cut-offs |
|---|---|---|
cutoff_10pct |
350 ms | 9.4 % |
cutoff_5pct |
577 ms | 4.8 % |
cutoff_2pct |
875 ms | 2.0 % |
Session
Open, configure and close the session.
session.startyou sendThe first frame. Authenticates and configures a stream session.
Rejections close the socket with 1008 and one of these reasons:
invalid api key, unknown mode '<mode>', a language error
(bad language code(s) [...], language(s) not supported by this ASR: [...],
unknown language mode ...), or an end-of-turn error
(eot: give a profile or a latency_budget_ms, not both,
unknown eot profile '<name>'; published: [...],
eot.latency_budget_ms must be a positive number, got ...).
| Field | Type | Description | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
type required | "session.start" | |||||||||||||||||
api_key required | string | |||||||||||||||||
mode | "transactional_call" | "working_session" | "one_on_one" | "brainstorm" | "formal" | Default | ||||||||||||||||
language | one of: Pinned, Allowed set, Automatic | Default | ||||||||||||||||
eot | one of: By profile, By latency budget | Choose an end-of-turn operating point by name, or by the mean latency you can
afford (the closest published point wins; ties go to the faster one). Give at most
one; omit
English figures; every language's points are at | ||||||||||||||||
signals | array of "vad" | "eot" | "bargein" | VAD and end of turn are always on. Add | ||||||||||||||||
idle_timeout_s | number | Close after this many seconds without speech. 0 or less disables it. Default | ||||||||||||||||
inference_interval | number | Scoring step in seconds. Event times fall on this grid. Default | ||||||||||||||||
messages | array of object | Conversation so far, used to judge whether an utterance is complete. |
A voice agent that wants barge-in.
{
"type": "session.start",
"api_key": "baton_0123456789abcdef0123456789abcdef0123456789abcdef",
"eot": {"profile": "cutoff_5pct"},
"language": {"pin": "en"},
"signals": ["vad", "eot", "bargein"]
}
Pick the operating point closest to a 600 ms mean latency.
{
"type": "session.start",
"api_key": "baton_0123456789abcdef0123456789abcdef0123456789abcdef",
"eot": {"latency_budget_ms": 600}
}
session.startedyou receiveThe session is open. Start sending audio.
| Field | Type | Description |
|---|---|---|
type required | "session.started" | |
schema_version required | 2 | |
session_id required | string | |
tier required | string | The scoring tier: |
eot required | object | The operating point this session uses. |
models | object | Present when |
{
"type": "session.started",
"schema_version": 2,
"session_id": "3f7c2a9e5b1d4c8e9a0b6d2f4e8c1a7b",
"tier": "full",
"eot": {
"profile": "cutoff_5pct",
"mean_latency_ms": 577,
"cutoff_rate": 0.0482
},
"models": {"bargein": "v1"}
}
warning (idle_timeout)you receiveNo speech for a while. The session will close soon.
Sent before an idle close (idle_timeout_s, default 600 s). Speech resets the
timer. The close is code 1000 with reason idle_timeout.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "warning" | |
code required | "idle_timeout" | |
close_in_s required | number | Seconds until the idle close. |
{
"type": "warning",
"seq": 40,
"schema_version": 2,
"code": "idle_timeout",
"close_in_s": 30.0
}
session.endyou sendEnd the session. Baton replies with session.summary and closes with 1000.
| Field | Type | Description |
|---|---|---|
type required | "session.end" |
{"type": "session.end"}
session.summaryyou receiveTotals for the session. Sent in reply to session.end.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "session.summary" | |
audio_s required | number | Seconds of audio received. |
turns required | integer | Turn ends |
{
"type": "session.summary",
"seq": 41,
"schema_version": 2,
"audio_s": 312.4,
"turns": 18
}
Audio
16 kHz mono PCM16 in binary frames.
audioyou sendOne binary frame of audio.
A binary WebSocket frame: one leading byte (send 0; it identifies the source
and is reserved for multi-source sessions), then signed 16-bit little-endian PCM,
mono, 16 000 Hz. Any whole number of samples per frame; 80–100 ms (1280–1600
samples) is typical. Resample on the client: other rates are not accepted.
Binary frame. [u8 source = 0][PCM16LE mono 16 kHz samples...]
Speech and turns
When people speak, pause and finish.
vad.startyou receiveThe user started speaking.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "vad.start" | |
t required | number | Audio time in seconds since the session's first sample. |
speaker required | "user" |
{
"type": "vad.start",
"seq": 1,
"schema_version": 2,
"t": 0.1,
"speaker": "user"
}
vad.endyou receiveThe user stopped making speech sound. Not yet a turn end.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "vad.end" | |
t required | number | Audio time in seconds since the session's first sample. |
speaker required | "user" | |
duration required | number | Seconds since |
{
"type": "vad.end",
"seq": 2,
"schema_version": 2,
"t": 2.3,
"speaker": "user",
"duration": 2.2
}
turn.endyou receiveThe user has finished their turn. Your agent may respond.
Fired after the speaker goes quiet, either because Baton is confident the turn is
complete (by: score) or because the silence has gone on long enough
(by: timeout). latency_ms is the silence between vad.end and this event.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "turn.end" | |
t required | number | Audio time in seconds since the session's first sample. |
speaker required | "user" | |
p_eot required | number | End-of-turn score at this step. |
latency_ms required | integer | Silence between |
by required | "score" | "timeout" | |
profile required | string |
{
"type": "turn.end",
"seq": 3,
"schema_version": 2,
"t": 3.0,
"speaker": "user",
"p_eot": 0.81,
"latency_ms": 700,
"by": "score",
"profile": "cutoff_5pct"
}
turn.resumedyou receiveThe user started speaking again within 2 s of turn.end. Treat that turn end as cancelled.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "turn.resumed" | |
t required | number | Audio time in seconds since the session's first sample. |
speaker required | "user" | |
after_ms required | integer | Milliseconds since the cancelled |
{
"type": "turn.resumed",
"seq": 5,
"schema_version": 2,
"t": 3.9,
"speaker": "user",
"after_ms": 900
}
Barge-in
Speech over your agent: a real interruption, or a backchannel.
agent.speechyou sendTell Baton when your agent starts and stops talking.
Needed for barge-in. While the agent is speaking, a user speech onset becomes a
bargein.onset and is then classified. Sending phase: end drops an onset that has
not been classified yet. Timestamped on arrival against the audio received so far.
Ignored unless the session subscribed to bargein.
| Field | Type | Description |
|---|---|---|
type required | "agent.speech" | |
phase required | "start" | "end" | |
utterance_id | string | Your id for the agent utterance. Echoed on barge-in events. |
{
"type": "agent.speech",
"phase": "start",
"utterance_id": "sp_1"
}
{
"type": "agent.speech",
"phase": "end",
"utterance_id": "sp_1"
}
bargein.onsetyou receiveThe user started speaking while your agent was talking.
Sent at the speech onset. Classification follows in bargein.class.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "bargein.onset" | |
t required | number | Audio time in seconds since the session's first sample. |
agent_speaking required | true | |
utterance_id | string |
{
"type": "bargein.onset",
"seq": 7,
"schema_version": 2,
"t": 12.1,
"agent_speaking": true,
"utterance_id": "sp_1"
}
bargein.classyou receiveWhether the interruption claims the floor or is a backchannel.
claim means the user wants to talk: stop or yield. backchannel ("mm-hm",
"right") means keep talking. Decided as soon as Baton is confident, typically a few
hundred milliseconds into the interruption. p_claim is null when the user
stopped too soon to judge; that case is reported as a backchannel.
| Field | Type | Description |
|---|---|---|
seq required | integer | Event sequence number. Starts at 1 and increases by one per event. |
schema_version required | 2 | |
type required | "bargein.class" | |
t required | number | Audio time in seconds since the session's first sample. |
class required | "claim" | "backchannel" | |
p_claim required | number or null | |
decided_after_ms required | integer | Milliseconds from the onset to this decision. |
utterance_id | string |
{
"type": "bargein.class",
"seq": 8,
"schema_version": 2,
"t": 12.4,
"class": "claim",
"p_claim": 0.9,
"decided_after_ms": 300,
"utterance_id": "sp_1"
}