Skip to content

feat: keep a model running with deployments - #2386

Draft
AlexCheema wants to merge 4 commits into
mainfrom
feat/keep-models-running
Draft

AlexCheema wants to merge 4 commits into
mainfrom
feat/keep-models-running

Conversation

@AlexCheema

Copy link
Copy Markdown
Contributor

Problem

An instance lives only as long as the nodes it was placed on. When one of them leaves (a crash, a restart, a network split), the master deletes the instance, and nothing places the model again: requests for it get 404 until someone notices and places it by hand. A cluster meant to serve a model needs a person for every node fault.

What this does

Adds deployments: a request to keep a model running. While a deployment exists, the master places an instance of its model whenever the cluster has none, with the same placement /place_instance uses.

POST   /deployments                 {"model_id": ..., "sharding": ..., "instance_meta": ..., "min_nodes": ...}
GET    /deployments                 each deployment with a status: serving | starting | placing | cant_place
DELETE /deployments/{deployment_id} stop, and delete the instance the deployment placed

It is a layer on top of instances, which are unchanged:

  • The keeper (src/exo/master/keeper.py) only ever adds an instance, and only when the cluster has no instance of the model at all. Any instance of the model counts, however it was placed. Deciding that an instance is broken stays where it is today: a node going silent, runners that keep failing, a user deleting it.
  • Deleting a deployment deletes only the instance the keeper placed for it, exactly as DELETE /instance does. An instance of the model placed some other way is left running.
  • Nothing changes unless a deployment exists.

Behaviour

  • Placement: once a second, at most one placement per second across all deployments, so each placement is made against a state that already holds the previous one.
  • Backoff: an instance lost before all its runners were ready counts as a failed placement: the next one waits 10 s, doubling up to 60 s. One lost after it was ready (its node died, someone deleted it) is placed again at once, and the backoff resets.
  • Nothing fits: the reason is recorded on the deployment (placementError, status cant_place) and reported only when it changes. It is tried again as soon as the nodes placement can use, or the links between them, change, and every 30 s otherwise.
    • A node rejoining arrives in pieces: it shows up, then reports memory and backends, then its links follow seconds later. The keeper tries again as each piece arrives.
  • New master: waits 15 s before placing anything, so nodes can reconnect and report the instances they already run.
  • Crashing placement: if placement raises something unexpected for one deployment, that deployment records the error and retries later. It doesn't stop the keeper keeping the others.

Races handled

The master's state trails the events it sends: they reach the state after a round trip through the event router.

  • Waiting for its own placement: the keeper remembers the instance it placed until the state shows it, and doesn't place again for 30 s meanwhile.
  • Deletion during a placement: a deployment's deletion and the keeper's events are emitted under one lock. The deletion also deletes a placement still on its way, which the cluster applies after the instance's creation.
  • Duplicate deployments: two requests to deploy the same model can both pass the API's 409 check. apply keeps the first, so the state never holds two deployments of one model.

Tests

  • test_keeper.py (37 tests, 56 cases with seeds): the keeper as a pure function of State, stepping it at chosen times and applying its events as the cluster does. Covers:
    • grace period, placement options, one placement per step, never touching existing instances;
    • in-flight placements, an instance lost between two steps, re-placement after node death;
    • the backoff sequence and its reset;
    • unplaceable deployments: changing reasons, nodes joining in pieces, memory freeing up, a crashing placement;
    • deletion, including a placement still on its way and double deletes;
    • status, and memory not growing.
    • Two randomized simulations:
      • 4 h of random instance and node loss on a 5-node cluster;
      • 20 seeds × 2 h of deployments created and deleted while the state lags the events by up to 10 s.
        Throughout, no model ever has two instances; once things settle, exactly the deployed models have one.
    • Every rule was mutation-checked: removing any one of them fails at least one test.
  • test_master_deployments.py: a real Master, with its events going round through a stand-in router. Deploy, and an instance is placed; delete the instance, and another is placed; delete the deployment, and the instance goes and nothing comes back.
  • test_apply_deployments.py: apply, one per model, events for deleted deployments, serialization.
  • test_deployments_api.py: the three endpoints, 409, 404.

On hardware

Two Mac Studios (M3 Ultra, 96 GB) kept Llama-3.2-1B (one node) and Qwen3.8-27B (pipeline over both) running through deployments, under continuous load (12 clients). Faults were injected every 1–2.5 minutes: node crashes, master crashes, freezes of 5–45 s, runner kills, network partitions. The test harness never placed an instance itself.

Run Duration Requests Faults Requests with neither a result nor an error (own node not faulted)
1 35 min 2,858 15 0
2 3 h 16,174 68 0
  • When the model's instance survived a fault: it served a new request within 0–26 s of the fault ending.
  • When the instance was lost with its node: the keeper placed it again, and the model served within 24–40 s after a node crash, and 46 s median (63 s worst) after a long freeze. This includes rejoining, placing and loading a 27B model.
  • Retrying on node changes: found on this hardware. With only a 30 s retry, a rejoining node's model took 47 s to serve again; with it, 27 s.

🤖 Generated with Claude Code

AlexCheema and others added 4 commits October 3, 2026 14:01
A deployment asks the cluster to keep one instance of a model running. Once a
second the master's keeper checks each deployment, and when the cluster has no
instance of its model it places one with the same placement /place_instance
uses. Instances are unchanged: the keeper only adds an instance when none of
the model exists, and deletes only the one it placed when its deployment is
deleted.

- POST /deployments, GET /deployments (with a status), DELETE /deployments/{id}
- An instance lost before it was ready backs off the next placement
  (10 s doubling to 60 s); one lost after it was ready is placed again at once
- A placement that fits nowhere is retried every 30 s and its reason recorded
- A new master waits 15 s before placing, for nodes to report their instances
- The keeper and a deployment's deletion emit events under one lock, and the
  deletion also deletes a placement still on its way to the state
- apply keeps one deployment per model even when two requests race

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…es change

A model spread over two nodes that lost one waited up to 30 s after the node
came back before the keeper tried again. Placement is too costly to poll
faster (88 ms per attempt on 8 fully connected nodes), so retry when the set
of nodes placement can use changes: connected, with memory and backends
reported. A joining node appears in the topology a moment before it reports
its memory; retrying on the topology alone would fail and then wait 30 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On hardware, a node that rejoined reported its memory and backends at once,
but its links to the other nodes arrived 7 and 12 s later. The retry made on
the memory report found no cycle and the next waited 30 s. The keeper now
compares the usable nodes and the links between them (and whether each is
RDMA), so each piece of a joining node is a reason to try again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Only ValueError (no placement fits) was handled. Anything else placement
raised for one deployment escaped step() every second: it was logged each
time, and deployments after it in id order were never placed. Record it as
that deployment's placement error, try again later as for one that doesn't
fit, and log it with its traceback when it first happens.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant