Skip to content

ConvAI custom LLM: engine retries after ~4s with no content, and empty deltas don't stop the clock #1005

Description

@ubmids

What I'm doing

Serving a custom LLM for a ConvAI agent: an OpenAI-compatible
POST /v1/chat/completions streaming SSE, backed by a self-hosted model rather
than a hosted API, so first-token latency varies a lot depending on whether the
weights are resident.

What happens

Conversations drop, and the pattern in my server log is always the same. Three
POSTs arrive about four seconds apart, none of them completes, then the client
gets Server error and the line goes down.

Each retry appears to cancel the request that was already in flight, so a turn
that takes longer than about four seconds can never finish, no matter how
healthy the socket is. It isn't a disconnect, it's a race the slow turn always
loses.

The part that took me a while

I already had a keepalive: an empty delta ({"delta": {}}) every 2 s over the
whole stream. The socket stays open, so from the connection's point of view
everything is fine, and the drops kept happening anyway.

The timer seems to watch for content, not for bytes. Empty deltas don't
reset it.

How I convinced myself

I evicted the model to force a cold start, then emitted a single space as the
first content delta before doing any real work:

first content:     0.78 s   (the space)
first real words:  9.39 s
conversation:      survived

Same cold start without that leading space drops at ~4 s as usual. So one byte
of content is the entire difference between a dropped call and a 9-second one.

For scale: a warm turn on my setup is about 2.2 s and a cold one is 10 s or
more, so before the workaround, whether a conversation survived was basically a
coin flip on whether the weights happened to be in memory.

Workaround

Emit content immediately, then stream normally:

def frames():
    yield chunk({"content": " "})   # stops the clock, spoken as nothing
    ...                             # everything real follows

A single space costs one frame and is spoken as silence.

What I'd like to know

  1. Is the ~4 s window intended, and is it configurable per agent? I couldn't
    find it documented anywhere.
  2. If it's fixed, could it be documented? Everyone pointing ConvAI at a custom
    LLM that isn't a fast hosted endpoint will hit this, and the symptom
    (Server error, dropped call) points nowhere near the cause.
  3. Could an empty delta, or an SSE comment, count as liveness? That's what a
    keepalive is for, and it currently reads as silence.

I've got logs of a dropped turn and a surviving one side by side if that helps.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions