Install Apple's command-line developer tools if needed:
xcode-select --installFrom the repository root:
make
./download_model.sh ds4f-q2
./ds4The same build supports M3 and M5 Macs. Hardware-specific fast paths are selected automatically; no environment variable is needed to enable them. Leave other GPU and memory-heavy applications idle when comparing performance.
| Memory | Starting point |
|---|---|
| 64 GB | Flash Q2 with --ssd-streaming |
| 96 GB | Flash Q2; leave room for the context and other applications |
| 128 GB | Flash Q2, or GLM 5.3 Flash Q2 with modest initial context |
| 256 GB | Flash Q4/MXFP4 or GLM 5.3 Flash Q4 |
| 512 GB | Larger models, including PRO Q2 |
Here, Flash means DeepSeek V4 Flash. V4.1 Flash has different memory requirements; see its model guide.
These are starting points, not guarantees that every context or session count will fit. GLM 5.3 Flash Q2 is about 90 GiB before runtime allocations. Stop other memory-heavy workloads before loading it resident.
./download_model.sh glm53-q2
./ds4 -m gguf/GLM-5.3-Flash-Q2.gguf --ctx 32768For a model larger than RAM, start with automatic cache sizing:
./ds4 --ssd-streamingSee SSD streaming before increasing the expert cache. Do not bypass the memory guard just to make an oversized resident model start.
Two 128 GB Macs can run a larger model fully resident with a 50/50 routed-expert split. Thunderbolt RDMA is the low-latency option; TCP is also supported. Follow tensor parallel setup. For more than two machines, use pipeline parallelism.
For DeepSeek V4 Flash and PRO, --power 70 trades throughput for lower
sustained GPU load. V4.1 and GLM currently require --power 100.