Repository navigation
running Qwen3.6-35B-A3B on a 8-16 GB Mac at 15 Token/s #1940
apirathaiya
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
I got Qwen3.6-35B-A3B running on a 16 GB M5 MacBook Air by keeping the routed experts on SSD and reading only the ones selected by the MoE router. The runtime also predicts which experts the next layer may need and starts loading them early.
The model uses around 1.7 GB of MLX memory. No experts are dropped, top-8 routing is unchanged, and seeded output is byte-identical to stock
mlx_lmon the same 4-bit checkpoint.Results on an M5 MacBook Air with 16 GB RAM:
I have released AiiStream Q3.6 as open source under Apache-2.0:
https://github.com/apirathaiya/aiistream-q3.6
The same streaming technique is designed to run with less than 16 GB of RAM, including 8 GB Macs. I have only tested the current release directly on 16 GB so far, so the 8 GB performance figures remain an estimate.
All reactions