Repository navigation
fix(worker): give up on an instance only after failed starts in a row - #2387
Open
AlexCheema wants to merge 1 commit into
Open
AlexCheema wants to merge 1 commit into
AlexCheema wants to merge 1 commit into
Conversation
The worker asks for an instance to be deleted once it has created a runner for it EXO_MAX_INSTANCE_RETRIES (5) times. The count was only cleared when the instance was deleted, so it also counted runners that had started and served, then died: a node restart, a killed process. A healthy instance whose runner had been restarted a few times over its life was deleted at the next restart. In a chaos test a pipeline instance was deleted 9 minutes after it was placed, after two of its runners were killed. Clear the count when this node's runner for the instance is ready, so it only counts attempts that never got there. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
A worker asks for an instance to be deleted once it has created a runner for it
EXO_MAX_INSTANCE_RETRIES(5) times. That limit is meant for an instance that can't start. But the count was only cleared when the instance was deleted, so it also counted runners that had started, served, and later died: a killed process, a peer's runner failing, a node restart. Over its life, a healthy instance used up its starts one restart at a time, and the next restart deleted it.In a chaos test, a 2-node pipeline instance was deleted 9 minutes after it was placed, after two of its runners were killed. A pipeline instance recreates runners on every node when one dies, sometimes with a failed attempt along the way, so a few kills use up 5 starts.
Fix
Clear an instance's count when this node's runner for it is ready. The count now means "starts in a row that never got there".
Tests
test_instance_retries.pydrives the worker's real event loop:The first test fails without the fix.
On hardware: one Mac Studio (M3 Ultra) running Llama-3.2-1B. Six times, wait for its runner to be ready, then kill the runner process.
🤖 Generated with Claude Code