Skip to content

fix(worker): give up on an instance only after failed starts in a row - #2387

Open
AlexCheema wants to merge 1 commit into
mainfrom
fix/instance-retries-count-failed-starts
Open

AlexCheema wants to merge 1 commit into
mainfrom
fix/instance-retries-count-failed-starts

Conversation

@AlexCheema

Copy link
Copy Markdown
Contributor

Problem

A worker asks for an instance to be deleted once it has created a runner for it EXO_MAX_INSTANCE_RETRIES (5) times. That limit is meant for an instance that can't start. But the count was only cleared when the instance was deleted, so it also counted runners that had started, served, and later died: a killed process, a peer's runner failing, a node restart. Over its life, a healthy instance used up its starts one restart at a time, and the next restart deleted it.

In a chaos test, a 2-node pipeline instance was deleted 9 minutes after it was placed, after two of its runners were killed. A pipeline instance recreates runners on every node when one dies, sometimes with a failed attempt along the way, so a few kills use up 5 starts.

Fix

Clear an instance's count when this node's runner for it is ready. The count now means "starts in a row that never got there".

Tests

  • test_instance_retries.py drives the worker's real event loop:

    • a runner that starts and is killed 10 times never uses up the retries;
    • 5 failed starts in a row still add up to the limit;
    • another node's runner being ready doesn't count.

    The first test fails without the fix.

  • On hardware: one Mac Studio (M3 Ultra) running Llama-3.2-1B. Six times, wait for its runner to be ready, then kill the runner process.

Instance after kills 1–5 Kill 6 After
main still running instance already deleted 0 instances, 1 deletion request
this PR still running still running 1 instance, ready, 0 deletion requests

🤖 Generated with Claude Code

The worker asks for an instance to be deleted once it has created a runner
for it EXO_MAX_INSTANCE_RETRIES (5) times. The count was only cleared when
the instance was deleted, so it also counted runners that had started and
served, then died: a node restart, a killed process. A healthy instance whose
runner had been restarted a few times over its life was deleted at the next
restart. In a chaos test a pipeline instance was deleted 9 minutes after it
was placed, after two of its runners were killed.

Clear the count when this node's runner for the instance is ready, so it only
counts attempts that never got there.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant