Skip to content

fix(worker): ask once to delete an instance it gave up on - #2389

Open
AlexCheema wants to merge 1 commit into
mainfrom
fix/worker-asks-once-to-delete
Open

AlexCheema wants to merge 1 commit into
mainfrom
fix/worker-asks-once-to-delete

Conversation

@AlexCheema

Copy link
Copy Markdown
Contributor

Problem

When a worker runs out of start attempts for an instance (EXO_MAX_INSTANCE_RETRIES), it sends the master DeleteInstance and skips the instance. Its planning loop comes back to the instance every 0.1 s until the deletion arrives, and sends the request again each time. While the master is slow to act on it (busy, frozen, or changing), the worker floods the network with the same command, about 10 per second. In one chaos run a node sent 227 in 24 s.

Fix

Ask once. If the instance is still there 10 s later, ask again, since the request may have been lost, e.g. to a master that was briefly replaced.

Tests

  • test_deletion_request.py: runs the worker's planning loop for 1.5 s on an instance past its retry limit, with the master never deleting it. With this PR the worker asks once; on main it asked 14 times.
  • On hardware: two Mac Studios (M3 Ultra) with Llama-3.2-1B on the node that isn't master. Kill its runner 5 times (waiting for it to be ready each time) so the worker gives up on the instance, then freeze the master for 25 s right after the last kill.
Deletion requests from the worker
main 96
this PR 2 (the first, and one retry 10 s later)

During the freeze, the other node briefly became master. It got the requests but didn't have the instance, and the old master then returned with it. The retry is what still gets the instance deleted afterwards.

🤖 Generated with Claude Code

When an instance runs out of start attempts, the worker sends DeleteInstance
and skips it. Its planning loop comes back to the instance every 0.1 s until
the deletion arrives, and sent the request again each time: one node sent 227
in 24 s while the master was slow to act on the first.

Ask once, and again only if the instance is still there 10 s later, in case
the request was lost (e.g. while the master changed).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant