Skip to content

fix: align Linux MLX CPU dependencies - #2385

Open
lorenzozanee wants to merge 1 commit into
exo-explore:mainfrom
lorenzozanee:fix/linux-mlx-cpu-install
Open

lorenzozanee wants to merge 1 commit into
exo-explore:mainfrom
lorenzozanee:fix/linux-mlx-cpu-install

Conversation

@lorenzozanee

Copy link
Copy Markdown

Motivation

The Linux CPU setup can fail when building miniaudio without Python headers. After that is resolved, the current dependency sources install mlx 0.32.0 from a custom fork with mlx-cpu 0.31.2, which produces an import-time ABI error.

Fixes #2382

Changes

  • Add python3-dev and build-essential to the Debian setup command.
  • Pin Linux mlx-cpu to 0.32.0 and limit the custom Linux mlx wheels to the CUDA extras, allowing the CPU extra to use the matching PyPI packages.
  • Clarify that Linux CPU utilization and large-model performance depend on the model operations and backend.
  • Add a regression test for the CPU package version and source markers.

Why It Works

The CPU extra now resolves mlx and mlx-cpu from PyPI at the same version, while the CUDA extras retain their custom package sources. The documented Debian prerequisites include Python headers for native dependency builds. The change does not establish or promise multi-core inference scaling.

Test Plan

Manual Testing

  • An isolated Python 3.13.14 environment installed the PyPI mlx==0.32.0 and mlx-cpu==0.32.0 wheels; mlx.core imported and evaluated a CPU array.
  • Full Linux project installation and 9B inference on the reported 32-thread system were not tested.

Automated Testing

  • python tests/test_linux_mlx_cpu_dependency_contract.py
  • ruff check tests/test_linux_mlx_cpu_dependency_contract.py
  • ruff format --check tests/test_linux_mlx_cpu_dependency_contract.py
  • git diff --check

uv.lock could not be regenerated because the public custom wheel endpoint timed out. uv lock --check reports that the lock needs updating, so locked project sync remains unverified.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] 32 threads, 1 core working: Linux CPU inference is pinned to a single thread (after 5 broken install steps just to get there)

1 participant