Baton listens to a live conversation and sends your voice agent typed events over one WebSocket:
- End of turn.
turn.endfires when a speaker has finished, not merely paused.vad.startandvad.endmark each stretch of speech. - Barge-in. When someone talks over your agent,
bargein.classsays whether they are claiming the floor or just saying "mm-hm". - Per-speaker transcripts. For several people mixed on one channel, Baton separates the voices (up to 8) and transcribes each one live.
- Language per voice. Baton identifies each voice's language and switches its transcription within the languages you allow.
| Stream | Room | |
|---|---|---|
| Endpoint | /v2/stream |
/v2/room |
| Who is speaking | One person, typically your agent's caller | Several people mixed on one channel |
| End of turn | Yes | Yes, for the room as a whole |
| Barge-in | Yes | Yes |
| Per-speaker transcripts | No | Yes |
| Language per voice | No | Yes |
A room is a stream plus transcripts and languages: the same session options and the same events, with speaker: "room".
How it works
One caller asks a question, pauses mid-sentence, finishes, and later says "mm-hm" while your agent answers. Baton's end-of-turn confidence stays low through the pause, rises once the caller has finished, and turn.end fires when it is confident. When the caller speaks over your agent, its barge-in confidence decides between a real interruption and a backchannel:
The same exchange as it arrives on the socket:
// the caller starts speaking
{"type": "vad.start", "seq": 1, "t": 0.3, "speaker": "user"}
// a pause mid-sentence: not confident the turn is over
{"type": "vad.end", "seq": 2, "t": 1.5, "speaker": "user",
"duration": 1.2}
{"type": "vad.start", "seq": 3, "t": 1.8, "speaker": "user"}
{"type": "vad.end", "seq": 4, "t": 2.7, "speaker": "user",
"duration": 0.9}
// 700 ms later Baton is confident: respond now
{"type": "turn.end", "seq": 5, "t": 3.4, "speaker": "user",
"p_eot": 0.81, "latency_ms": 700, "by": "score",
"profile": "cutoff_5pct"}
// at 3.7 s you send agent.speech, phase start, utterance sp_1
// the caller speaks over your agent
{"type": "vad.start", "seq": 6, "t": 5.0, "speaker": "user"}
{"type": "bargein.onset", "seq": 7, "t": 5.0,
"agent_speaking": true, "utterance_id": "sp_1"}
// it was a backchannel: keep talking
{"type": "bargein.class", "seq": 8, "t": 5.3,
"class": "backchannel", "p_claim": 0.12,
"decided_after_ms": 300, "utterance_id": "sp_1"}
{"type": "vad.end", "seq": 9, "t": 5.35, "speaker": "user",
"duration": 0.35}
// at 6.6 s you send agent.speech, phase end, utterance sp_1
Every event also carries "schema_version": 2. t is seconds of audio you have sent. In a room, these events carry "speaker": "room", and per-speaker transcripts arrive alongside them.
Baton is in private beta. Request a key at hello@joinin.ai. The v2 API can still change; ignore fields and events you don't recognise.
Pick an endpoint
| Endpoint | For | You get | Latency at 5 % false cut-offs (English) |
|---|---|---|---|
/v2/stream |
One speaker: a voice agent and its caller. Can use the conversation so far. | VAD, end of turn, barge-in | 577 ms |
/v2/room |
Several people mixed on one channel. | VAD, end of turn, barge-in, per-speaker transcripts, language ID | 726 ms |
Latency is the mean delay from the end of speech to turn.end on Baton's benchmark, at the cutoff_5pct operating point. Every endpoint takes the same session options; see End of turn for the faster and more careful points.
Servers
| Environment | Host | Notes |
|---|---|---|
| Production | baton.joinin.ai |
The /v1 streams today. The /v2 endpoints are coming; use the test server meanwhile. |
| Test | baton-test.joinin.ai |
Every endpoint, including /v2/room. For integration testing; may be briefly unavailable. |