Repository navigation
fix(worker): ask once to delete an instance it gave up on - #2389
Open
AlexCheema wants to merge 1 commit into
Open
AlexCheema wants to merge 1 commit into
AlexCheema wants to merge 1 commit into
Conversation
When an instance runs out of start attempts, the worker sends DeleteInstance and skips it. Its planning loop comes back to the instance every 0.1 s until the deletion arrives, and sent the request again each time: one node sent 227 in 24 s while the master was slow to act on the first. Ask once, and again only if the instance is still there 10 s later, in case the request was lost (e.g. while the master changed). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
When a worker runs out of start attempts for an instance (
EXO_MAX_INSTANCE_RETRIES), it sends the masterDeleteInstanceand skips the instance. Its planning loop comes back to the instance every 0.1 s until the deletion arrives, and sends the request again each time. While the master is slow to act on it (busy, frozen, or changing), the worker floods the network with the same command, about 10 per second. In one chaos run a node sent 227 in 24 s.Fix
Ask once. If the instance is still there 10 s later, ask again, since the request may have been lost, e.g. to a master that was briefly replaced.
Tests
test_deletion_request.py: runs the worker's planning loop for 1.5 s on an instance past its retry limit, with the master never deleting it. With this PR the worker asks once; on main it asked 14 times.During the freeze, the other node briefly became master. It got the requests but didn't have the instance, and the old master then returned with it. The retry is what still gets the instance deleted afterwards.
🤖 Generated with Claude Code