Skip to content

fix(mlx): wait for the in-flight decode step before a prompt eval (fixes rare tensor-parallel deadlock) - #2393

Open
AlexCheema wants to merge 1 commit into
mainfrom
fix/no-prefill-during-decode-step
Open

AlexCheema wants to merge 1 commit into
mainfrom
fix/no-prefill-during-decode-step

Conversation

@AlexCheema

@AlexCheema AlexCheema commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Problem

A tensor-parallel model over RDMA (MlxJaccl, fast synch on) still deadlocks now and then under concurrent load, even with #2391. It happened about once an hour when another model's runner shared one of its Macs, and within minutes in chaos testing, where the cluster moves models around. The stuck-runner watchdog then kills the runner after 90 s and every request in flight fails.

What the stuck rank looks like (stack samples, lldb on the live process, a Metal System Trace):

  • Its communication stream spins in mlx::core::Fence::wait, waiting for its GPU to signal fence value 128. The update to 128 has been encoded (the fence's update count is 128), but the buffer holds 127.
  • The GPU side is waiting on the communication stream's fence, which already holds the value it wants (127).
  • The process's GPU work is one execution that never ends: a fast-synch spin wait, according to the trace.
  • Its main thread waits in eval_impl, having committed everything.
  • The other rank waits in jaccl::MeshImpl::all_reduce.

Cause

Two evals of the tensor-parallel model overlap:

  • mlx-lm's BatchGenerator._next async-evaluates the batch's next decode step, then processes new prompts straight away;
  • exo prefills a new request while the batch's last decode step is still running.

Each eval queues its collectives (128 for this 64-layer model) on the one communication stream, and its GPU work spins waiting for them (fast synch). The prompt eval's committed command buffers start spinning while the decode step's last kernels still have to run. MLX's buffers are untracked, so nothing orders those kernels against the spinning ones.

The two waits form a cycle:

  • the spinning kernels wait for the communication stream;
  • the communication stream waits for the decode step's last fence update;
  • that update is behind a kernel that doesn't get to run.

With the GPU to itself this rarely bites. With another process's work on the same GPU it happens about once an hour.

Fix

Don't let a prompt eval overlap a decode step. Before exo's prefill, and before the batch processes new prompt tokens, synchronize mlx-lm's generation stream. A new request waits for at most one decode step before its prefill starts; steady decoding is unchanged.

Evidence

Reproduction without exo (mlx-lm only: sharded_load + BatchGenerator over JACCL on two Mac minis, M4 Pro 64 GB, Thunderbolt 5 RDMA, MLX main 0e3ff364, fast synch on, a second mlx-lm process generating on the same Macs):

Sequences Hangs
As mlx-lm does it ~430 2
Prompt evals wait for the decode step ~620 0

exo (Qwen3.8-27B tensor parallel over RDMA on 2 Mac minis, Llama-3.2-1B on one of them, 6 + 4 concurrent clients):

Time Qwen requests failed Stuck-runner restarts
Before (#2391 applied) 100 min (10 runs) 12 2
This PR 6 × 30 min, two pairs of minis 0 of 2,364 0

No cost measured elsewhere.

Before This PR
Qwen tensor parallel alone, 6 clients, requests per 10 min 186 186, 192 (0 stalls)
Qwen tensor parallel, single requests, generation tok/s 26.0, 25.9 25.9, 25.9
Qwen tensor parallel, single requests, prompt tok/s 196.1, 196.5 196.1, 196.1
Llama-3.2-1B on one node, generation tok/s 266.4, 266.4 267.4, 268.0

Single requests: 731-token prompt, 256 tokens out, medians of 20, two alternating rounds.

Sharing a Mac with Llama, the Qwen instance completed 132 requests per 10 min, against 138 in the earlier runs that didn't stall (within run-to-run variation).

Test plan

  • uv run basedpyright (0 errors), uv run ruff check, ruff format
  • uv run pytest src -m "not slow" (CI's command): 471 passed, 3 skipped
  • New tests: the batch's prompt processing waits for the decode step first (not when there is nothing to process), and the patch installs it
  • Hardware: reproduction and A/B above

🤖 Generated with Claude Code

mlx-lm's BatchGenerator async-evaluates the batch's next decode step and then
processes new prompts, and exo prefills a new request while the batch's last
step is still running. For a tensor-parallel model with MLX_METAL_FAST_SYNCH,
both evals queue their collectives on the one communication stream, while
their GPU work spins waiting for it. When the prompt eval's spinning GPU work
held the GPU while the step's last kernels still had to run (more likely with
another model's runner on the same Mac), neither could finish: one rank's
communication stream waited in Fence::wait for a fence update that its GPU
never ran, and the other rank waited in all_reduce.

Synchronize mlx-lm's generation stream before exo's prefill and before the
batch processes new prompt tokens, so that no prompt eval overlaps a decode
step. A request now waits for at most one decode step before its prefill;
steady decoding is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant