What I'm doing
Serving a custom LLM for a ConvAI agent: an OpenAI-compatible
POST /v1/chat/completions streaming SSE, backed by a self-hosted model rather
than a hosted API, so first-token latency varies a lot depending on whether the
weights are resident.
What happens
Conversations drop, and the pattern in my server log is always the same. Three
POSTs arrive about four seconds apart, none of them completes, then the client
gets Server error and the line goes down.
Each retry appears to cancel the request that was already in flight, so a turn
that takes longer than about four seconds can never finish, no matter how
healthy the socket is. It isn't a disconnect, it's a race the slow turn always
loses.
The part that took me a while
I already had a keepalive: an empty delta ({"delta": {}}) every 2 s over the
whole stream. The socket stays open, so from the connection's point of view
everything is fine, and the drops kept happening anyway.
The timer seems to watch for content, not for bytes. Empty deltas don't
reset it.
How I convinced myself
I evicted the model to force a cold start, then emitted a single space as the
first content delta before doing any real work:
first content: 0.78 s (the space)
first real words: 9.39 s
conversation: survived
Same cold start without that leading space drops at ~4 s as usual. So one byte
of content is the entire difference between a dropped call and a 9-second one.
For scale: a warm turn on my setup is about 2.2 s and a cold one is 10 s or
more, so before the workaround, whether a conversation survived was basically a
coin flip on whether the weights happened to be in memory.
Workaround
Emit content immediately, then stream normally:
def frames():
yield chunk({"content": " "}) # stops the clock, spoken as nothing
... # everything real follows
A single space costs one frame and is spoken as silence.
What I'd like to know
- Is the ~4 s window intended, and is it configurable per agent? I couldn't
find it documented anywhere.
- If it's fixed, could it be documented? Everyone pointing ConvAI at a custom
LLM that isn't a fast hosted endpoint will hit this, and the symptom
(Server error, dropped call) points nowhere near the cause.
- Could an empty delta, or an SSE comment, count as liveness? That's what a
keepalive is for, and it currently reads as silence.
I've got logs of a dropped turn and a surviving one side by side if that helps.
What I'm doing
Serving a custom LLM for a ConvAI agent: an OpenAI-compatible
POST /v1/chat/completionsstreaming SSE, backed by a self-hosted model ratherthan a hosted API, so first-token latency varies a lot depending on whether the
weights are resident.
What happens
Conversations drop, and the pattern in my server log is always the same. Three
POSTs arrive about four seconds apart, none of them completes, then the client
gets
Server errorand the line goes down.Each retry appears to cancel the request that was already in flight, so a turn
that takes longer than about four seconds can never finish, no matter how
healthy the socket is. It isn't a disconnect, it's a race the slow turn always
loses.
The part that took me a while
I already had a keepalive: an empty delta (
{"delta": {}}) every 2 s over thewhole stream. The socket stays open, so from the connection's point of view
everything is fine, and the drops kept happening anyway.
The timer seems to watch for content, not for bytes. Empty deltas don't
reset it.
How I convinced myself
I evicted the model to force a cold start, then emitted a single space as the
first content delta before doing any real work:
Same cold start without that leading space drops at ~4 s as usual. So one byte
of content is the entire difference between a dropped call and a 9-second one.
For scale: a warm turn on my setup is about 2.2 s and a cold one is 10 s or
more, so before the workaround, whether a conversation survived was basically a
coin flip on whether the weights happened to be in memory.
Workaround
Emit content immediately, then stream normally:
A single space costs one frame and is spoken as silence.
What I'd like to know
find it documented anywhere.
LLM that isn't a fast hosted endpoint will hit this, and the symptom
(
Server error, dropped call) points nowhere near the cause.keepalive is for, and it currently reads as silence.
I've got logs of a dropped turn and a surviving one side by side if that helps.