Repository navigation
DeepSeek-V4-Pro-0813 TP4 + DSpark on four M3 Ultra (mlx.launch) #1728
Replies: 4 comments
|
Recipe and scripts (convert / shard_mtp / BatchGenerator smoke / OpenAI serve). No converted weights — rebuild from https://huggingface.co/guruswami-ai/deepseek-v4-pro-0813-mlx-tp4-recipe |
Interim: singleton serve, second generate, prefix trimThe OpenAI endpoint ( The JACCL Recv/Send -12 that killed later generates was an undrained Stock Recipe (scripts, no weights): https://huggingface.co/guruswami-ai/deepseek-v4-pro-0813-mlx-tp4-recipe |
Follow-up, 15 August 2026Same stack (world=4, gather, ContextOne process, cap 131 072, 32 new tokens, greedy. Decode stays ~12 tok/s. Prefill is what dies with length.
129 k is usable if you can wait twelve minutes for the first token. It is not comfortable. Prefix cache and streamingPoolingCache still cannot large-trim. Chasing a better trim is the wrong next step. What works is an append-only tape of post-template token ids: a dsh-style follow-up logs SSE is no longer a replay of a finished completion. Ranks stay in lockstep on each token. One framing bug blocked stock urllib: Still required
Asks from the first post still stand. Happy to paste scripts. Not claiming this is upstream-ready. |
|
Parking Pro-0813 TP4 as a daily driver (17 August 2026). The load path and short-prompt benches in this thread still hold. Long-lived JACCL dies independently of the model: ml-explore/mlx#4319 and Apple FB24371487. Full note on mlx discussion 4247. |
Uh oh!
There was an error while loading. Please reload this page.
DeepSeek-V4-Pro-0813 TP4 + DSpark on four M3 Ultra (mlx.launch)
Report from a four-node M3 Ultra 512 GB mesh (Thunderbolt 5, JACCL). Full write-up with tables: we can paste more detail or open a gist if useful.
Result: one interactive stream of Pro-0813 (1.65 T / 49 B active, official MXFP4 experts + MXFP8 attn, MTP kept) at 15–23 tok/s decode with the checkpoint's own DSpark drafter, out to 32 768 prompt tokens. An OpenAI-compatible server in front of
mlx.launchis enough for DeepSeek Harness (dsh) to run a real tool loop (failing unit tests → model edits → tests pass).This is the first trillion-class open-weight pack we have run on this mesh that is usable as an agent. Previous 2 T-class attempts either did not fit or collapsed into tokenizer boilerplate.
Why this is an mlx-lm issue
Stock mlx-lm 0.31.3 has no
deepseek_v4. We used the omlx 0.5.7 overlay (deepseek_v4+mlx_lm_mtp) as a library undermlx.launch, not the omlx product server.PRs already open: #1189, #1233. Until one lands, anyone reproducing this is on a patched tree.
mlx_lm.serverrefuses draft models in distributed mode. DSpark only arms inside the patchedGenerationBatch.next().stream_generate/generate_stepnever activate MTP. The working driver isBatchGenerator.Model.shard()does not shardself.mtp. The three DSpark stages need the same head/expert split as the backbone or TP4 desynchronises.Sampled TP4 without
mx.random.seedon every rank parked at 49/48/45/47 tokens and killed JACCL (Send failed … -12). Shared seed is mandatory.Numbers (14 August 2026, world=4, gather fallback, FAST_SYNCH=0)
17*23=→ 391Without DSpark the same math smoke is 5 tok/s. With it, 19 tok/s / 96% accept on the first 128-token run.
B>1 is backbone only (
_omlx_mtp_rowwise_unsupportedon DSpark). One DSpark user is faster than four backbone users.What blocked long context
Near 4 k prompt:
ValueError: Unsupported DSpark physical-ring GEMM layout. omlx'sdspark_ring_gemmis compiled for 64 heads. Pro TP4 has 32 heads/rank. Python enables that kernel once pooled KV exceedsindex_topk1024. Hiding the symbol forces the gather path inSparseCompressedAttention. That unlocked 4 k / 16 k / 32 k. A 32-head native kernel compiled but lowered accept (70% → 56%) and then tore JACCL at 16 k. Serving stays on gather.Instability
JACCL Recv/Send -12 on a later generate in the same
mlx.launchprocess is still common. First generate after load usually works. Second HTTP turn or secondBatchGeneratorsometimes kills a rank.mlx.launchthen reportsexit 0. After a crash, leftover ranks plus a hungibv_devinfoneed a full-fleet reboot and one freshmlx.distributed_config --auto-setup.Asks
deepseek_v4and keep MTP tensors loadable.mlx_lm.serverdistributed DSpark, or document that fused draft isBatchGeneratoronly.Happy to share scripts (
convert,shard_mtp,BatchGeneratorbench, OpenAI serve with an arm barrier). Not claiming this is upstream-ready. It works well enough that a person can sit in DeepSeek Harness and get code back. That is new for us at this scale on Apple Silicon.All reactions