Repository navigation
feat: keep a model running with deployments - #2386
Draft
AlexCheema wants to merge 4 commits into
Draft
AlexCheema wants to merge 4 commits into
AlexCheema wants to merge 4 commits into
Conversation
A deployment asks the cluster to keep one instance of a model running. Once a
second the master's keeper checks each deployment, and when the cluster has no
instance of its model it places one with the same placement /place_instance
uses. Instances are unchanged: the keeper only adds an instance when none of
the model exists, and deletes only the one it placed when its deployment is
deleted.
- POST /deployments, GET /deployments (with a status), DELETE /deployments/{id}
- An instance lost before it was ready backs off the next placement
(10 s doubling to 60 s); one lost after it was ready is placed again at once
- A placement that fits nowhere is retried every 30 s and its reason recorded
- A new master waits 15 s before placing, for nodes to report their instances
- The keeper and a deployment's deletion emit events under one lock, and the
deletion also deletes a placement still on its way to the state
- apply keeps one deployment per model even when two requests race
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…es change A model spread over two nodes that lost one waited up to 30 s after the node came back before the keeper tried again. Placement is too costly to poll faster (88 ms per attempt on 8 fully connected nodes), so retry when the set of nodes placement can use changes: connected, with memory and backends reported. A joining node appears in the topology a moment before it reports its memory; retrying on the topology alone would fail and then wait 30 s. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On hardware, a node that rejoined reported its memory and backends at once, but its links to the other nodes arrived 7 and 12 s later. The retry made on the memory report found no cycle and the next waited 30 s. The keeper now compares the usable nodes and the links between them (and whether each is RDMA), so each piece of a joining node is a reason to try again. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Only ValueError (no placement fits) was handled. Anything else placement raised for one deployment escaped step() every second: it was logged each time, and deployments after it in id order were never placed. Record it as that deployment's placement error, try again later as for one that doesn't fit, and log it with its traceback when it first happens. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
An instance lives only as long as the nodes it was placed on. When one of them leaves (a crash, a restart, a network split), the master deletes the instance, and nothing places the model again: requests for it get 404 until someone notices and places it by hand. A cluster meant to serve a model needs a person for every node fault.
What this does
Adds deployments: a request to keep a model running. While a deployment exists, the master places an instance of its model whenever the cluster has none, with the same placement
/place_instanceuses.It is a layer on top of instances, which are unchanged:
src/exo/master/keeper.py) only ever adds an instance, and only when the cluster has no instance of the model at all. Any instance of the model counts, however it was placed. Deciding that an instance is broken stays where it is today: a node going silent, runners that keep failing, a user deleting it.DELETE /instancedoes. An instance of the model placed some other way is left running.Behaviour
placementError, statuscant_place) and reported only when it changes. It is tried again as soon as the nodes placement can use, or the links between them, change, and every 30 s otherwise.Races handled
The master's state trails the events it sends: they reach the state after a round trip through the event router.
applykeeps the first, so the state never holds two deployments of one model.Tests
test_keeper.py(37 tests, 56 cases with seeds): the keeper as a pure function ofState, stepping it at chosen times and applying its events as the cluster does. Covers:Throughout, no model ever has two instances; once things settle, exactly the deployed models have one.
test_master_deployments.py: a realMaster, with its events going round through a stand-in router. Deploy, and an instance is placed; delete the instance, and another is placed; delete the deployment, and the instance goes and nothing comes back.test_apply_deployments.py: apply, one per model, events for deleted deployments, serialization.test_deployments_api.py: the three endpoints, 409, 404.On hardware
Two Mac Studios (M3 Ultra, 96 GB) kept Llama-3.2-1B (one node) and Qwen3.8-27B (pipeline over both) running through deployments, under continuous load (12 clients). Faults were injected every 1–2.5 minutes: node crashes, master crashes, freezes of 5–45 s, runner kills, network partitions. The test harness never placed an instance itself.
🤖 Generated with Claude Code