People pause mid-sentence all the time. Baton continuously judges whether the speaker is done, and sends turn.end only when it is confident they have finished.
How a turn ends
turn.end arrives after the speaker goes quiet, in one of two ways:
- Baton is confident the turn is complete (
by: "score").p_eotis that confidence. - The silence has gone on long enough that the turn ends anyway (
by: "timeout"). The limit depends on the operating point.
If the speaker starts again within 2 s of turn.end, Baton sends turn.resumed. Treat that turn end as cancelled.
Operating points
Each operating point is a published trade-off between speed and false cut-offs (pauses wrongly ended as turns), measured on Baton's end-of-turn benchmark.
| Operating point | /v2/stream |
/v2/room |
False cut-offs | Good for |
|---|---|---|---|---|
cutoff_10pct |
350 ms | 439 ms | ≈ 10 % | Quick back-and-forth: bookings, support calls |
cutoff_5pct |
577 ms | 726 ms | ≈ 5 % | Conversations and meetings |
cutoff_2pct |
875 ms | 1134 ms | ≈ 2 % | People who think aloud: interviews, brainstorms |
Latency is the mean delay from the end of speech to turn.end. English values; points for all 14 supported languages are published at GET /v2/eot/profiles.
Choosing one
Set eot in session.start, by name or by the latency you can afford. Without eot, a session uses cutoff_10pct.
{
"type": "session.start",
"api_key": "baton_...",
"eot": {"profile": "cutoff_5pct"}
}
{
"type": "session.start",
"api_key": "baton_...",
"eot": {"latency_budget_ms": 600}
}
A latency budget picks the published point with the closest mean latency; ties go to the faster one. With {"latency_budget_ms": 600}, /v2/stream gets cutoff_5pct (577 ms) and /v2/room gets cutoff_5pct (726 ms); with 500, a room gets cutoff_10pct (439 ms).
session.started reports the point in use:
{
"type": "session.started",
"schema_version": 2,
"session_id": "3f7c...",
"tier": "full",
"eot": {
"profile": "cutoff_5pct",
"mean_latency_ms": 577,
"cutoff_rate": 0.0482
}
}
Conversation context
On /v2/stream, pass the conversation so far in messages. Baton uses it to judge whether an utterance is complete: "Paris" after "Where are you flying?" is a full answer.
{
"type": "session.start",
"api_key": "baton_...",
"messages": [
{
"role": "assistant",
"content": "What city are you flying to?"
}
]
}