MLX Community Projects #654
Replies: 72 comments 21 replies
|
text generation: https://github.com/mzbac/mlx-moe-models |
|
An implementation of Reinforcement Learning algorithms in MLX based in the Implementations from CleanRL. Still WIP because it’s missing a benchmark and some other minor things, but the implementations work correctly. |
|
mlx-models. Currently supporting vision models by loading/converting from PyTorch checkpoints. Will later add support for text and audio models as well. |
|
Hi I would love to add chat-with-mlx. It is a Chat UI + RAG Implementation on MLX. I wIll add more features later on (more advanced RAG pipeline + multimodal) |
|
I have an example of training a simple language model using BitLinear instead of nn.Linear. It's a port of Karpathy's minGPT to MLX along with a custom implementation of a BitLinear module. https://github.com/adhulipa/mlx-mingpt I noticed this collection already has the far more meatier |
|
Transformer Lab https://github.com/transformerlab/transformerlab-app is an LLM research platform that allows you to run, train, perform RAG, and evaluate LLMs through a GUI. |
|
MLX RAG with GGUF Models: https://github.com/Jaykef/mlx-rag-gguf The code here builds on https://github.com/vegaluisjose/mlx-rag, it has been optimized to support RAG-based inferencing for .gguf models. I am using BAAI/bge-small-en for the embedding model, TinyLlama-1.1B-Chat-v1.0-GGUF as base model and the custom vector database script for indexing texts in a pdf file. Inference speeds can go up to ~413 tokens/sec for prompts and ~36 tokens/sec for generation on my 8G M2 Air. |
|
Vision: MLX3D A library for deep learning with 3D data using mlx. |
|
JSON schema decoding (allowing function calling, including an OpenAI-compatible server with tools) using MLX: https://github.com/otriscon/llm-structured-output |
|
Hello for text generation part, I'm happy to share with you that I've proposed and contributed to the integration of MLX with LibreChat.ai. So now you can use your local LLM powered by MLX through a fancy interface privately, enjoy! :D See LibreChat-AI/LibreChat#2580 If in the future the community proposes an API servers supporting also multimodality, transcription, image generation for example, I will add them into LibreChat ;) It could be great also to have and LLM API supporting /models endpoint and multiple models simultaneously :D |
|
Hello, mlx community, we are happy to share with you that we have contributed the first strong sub-4 bit LLM model zoo for MLX community.
The modern LLM families include Llama3/2, Phi-3, Mistral, 01-Yi, and Qwen. A mlx-style inference toolkit is also shared for the local web chatting.
We are an active team here, supporting the better low-bit community on the local platform. Enjoy! |
|
mlx_micrograd - mlx port of Karpathy's micrograd - a tiny scalar-valued autograd engine with a small PyTorch-like neural network library on top. Installationpip install mlx_microgradExample usageExample showing a number of possible supported operations: from mlx_micrograd.engine import Value
a = Value(-4.0)
b = Value(2.0)
c = a + b
d = a * b + b**3
c += c + 1
c += 1 + c + (-a)
d += d * 2 + (b + a).relu()
d += 3 * d + (b - a).relu()
e = c - d
f = e**2
g = f / 2.0
g += 10.0 / f
print(f'{g.data}') # prints array(24.7041, dtype=float32), the outcome of this forward pass
g.backward()
print(f'{a.grad}') # prints array(138.834, dtype=float32), i.e. the numerical value of dg/da
print(f'{b.grad}') # prints array(645.577, dtype=float32), i.e. the numerical value of dg/db
|
|
mlx-serve — native Zig inference server for Apple Silicon, no Python. Runs MLX-format models and exposes OpenAI-compatible and Anthropic-compatible HTTP APIs out of the box, so the same Highlights:
MIT, Apple-Silicon only. Site: ddalcu.github.io/mlx-serve |
|
atlas — measured-cost quantization for MLX: profiles per-block KL sensitivity, then solves (bit-width, group-size) allocation exactly for any RAM budget. Benchmarked vs uniform MLX and llama.cpp K-quants. https://github.com/Matth21/atlas |
|
I built something that cuts local model latency and footprint. Squish is an MLX-based local inference server for Apple Silicon, no VRAM, no CUDA. Measured against Ollama on an M3 16GB, 1.15x to 14.7x faster depending on how much your prompts repeat, plus a smaller memory footprint on top. Only tested on a 16GB M3 so far, looking for feedback and testers with more unified memory or a newer M series chip to see how it handles bigger models. |
|
A few more MLX tools from me, this time on the developer and diagnostics side (the two diffusion ones I posted earlier are already in the list):
Thanks for keeping this list going! |
|
PostTrainLLM — Mac-local post-training + eval factory (specialists, report cards, MLX packaging).
Happy to adjust the blurb if a shorter line is preferred. |
|
mlx-tsfm: time-series foundation models on Apple Silicon. The first release is an MLX port of TimesFM 3.0 (330M), checked against the PyTorch reference at about 1e-6 max abs error. https://github.com/rachittshah/mlx-tsfm |
|
MindCraft Studio — MLX-Powered AI Image Generator for Mac |
|
Isla Watson — an independent local AI companion research project from Scotland. I’ve just posted an architecture-level overview exploring MLX/Apple Silicon as a local compute fabric for a persistent companion: replaceable foundation models, persistent identity/state outside the model, specialist Mac agents, and distributed inference when larger models are needed. The project is currently at the MLX evaluation/prototyping stage, and I’d particularly value feedback from anyone working with distributed MLX or multi-Mac workloads. Full discussion: #4482 |
|
mlx.fast (https://mlx.fast, leaderboard at https://www.yukon.org/mlxfast) We built an open benchmark arena for MLX inference speed on Apple Silicon. The current task is faster prefill and serial decode for Qwen 3.8 Flash Next with speculative decoding, scored as a single composite ( The harness is open source. Clone it, run locally, submit: https://github.com/Layr-Labs/mlxfast-qwen38-125b-a6b-engine Suggested entry for Misc:
(The harness is built on MLX Swift, so happy to have this on the MLX Swift Community Projects thread instead, or both, whichever you prefer.) |
|
VeloxQuant-MLX: Shrinks the KV cache of any |
|
I'd like to share Open Decisions, an open-source Python package for typed decisions from local models, inspired by Jev. It has MLX-LM and MLX-VLM adapters; I tested Qwen3-VL-4B-Instruct in 4-bit on an M4 Pro. The package prefills the prompt, reads next-token logits for labels representing the allowed choices, and converts them into choices, yes/no scores, or joint boolean gates. No answer tokens are generated; inference still runs, and the probabilities are uncalibrated. The agent-routing example scores whether to delegate and whether to ask for missing context in one evaluation. Three recorded routing requests took 533 ms with an uncached policy, then 113 ms and 118 ms with the policy cached, with the model already loaded. These are individual observations, not a benchmark. There's also a recorded Tetris comparison with hosted Jev. The controller supplies structured board state, legal landings and projected outcomes, so the example has substantial help from code and uses no screenshots. The repo includes installation instructions, runnable examples, recordings and limitations. I'd welcome feedback on the MLX adapters and evaluating these decisions across models. |
|
two MLX projects from me: claude-code-local: runs Claude Code fully on-device. A native MLX server (built on nemotron-omni-mlx: pure MLX runtime for the vision and audio towers of NVIDIA's Nemotron 3 Nano Omni, which had no Apple Silicon runtime. It passes 23/23 parity tests against NVIDIA's PyTorch reference. 67.7 tok/s with an image and 147 tok/s with audio on an M5 Max, MIT licensed. I also added Muse Glimmer text model support to mlx-lm (ml-explore/mlx-lm#1710). Thanks to everyone who works on MLX, none of this would exist without it. |
|
laya-apple: a heterogeneous MLX GPU + Apple Neural Engine runtime for Laya typed-decision models. Neural Engine artifacts are used only after they pass a parity check against upstream Laya on the machine that runs them. On one M4 Max, under a mixed short + long workload, serving on the GPU and the ANE at the same time reached up to 4.57× the throughput of GPU-only serving (method and data). I'm collecting benchmarks from other Apple silicon: community matrix and write-up. Would you consider adding it to the Community Projects list? |
|
mlx-recipes — end-to-end ML recipes in pure MLX. https://github.com/nperumal/mlx-recipes I've been building a small repo of end-to-end recipes written in pure MLX — no NumPy or scikit-learn anywhere in the math path, enforced by a CI grep. Two recipes so far, both in a foundations tier aimed at people learning the framework rather than at benchmarking it: linear regression two ways (mx.grad with a hand-written update vs nn.Linear + mlx.optimizers), and logistic regression binary-then-softmax on Palmer Penguins. |
|
If you tune a model on a Mac, how do you know which version is serving now? Cohesix 1.2.0 keeps the comparison, serving check, and rollback tied to the same model change. MLX still runs locally on Apple Silicon; the guide also covers optional vMLX serving. The Mac rollout walkthrough is here: |


Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Let's collect some cool MLX integrations and community lead projects here for visibility!
If you have a project you would like to feature, leave a comment, and we will add it. If the project is build with MLX Swift, add it to the MLX Swift Community Project page.
Text Generation
Vision
Speech and Audio
Multi-modal
Misc
Educational
picoGPT.All reactions