Skip to content

decision model on the neural engine (measured) :3 - #56

Merged
huyedits merged 3 commits into
mainfrom
symbio-pet
Sep 26, 2026
Merged

huyedits merged 3 commits into
mainfrom
symbio-pet

Conversation

@huyedits

@huyedits huyedits commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Follow-up to #55.

Measured with macmon, which reads the power rails without sudo. Each phase repeats one operation for 6 s.

operation ANE CPU GPU
idle 0.007 W ~0.8 W ~0.07 W
OCR, .accurate, pinned to the ANE 0.52 W 5.2 W 0.07 W
OCR, .fast (CPU-only control) 0.002 W 6.2 W 0.11 W
NLContextualEmbedding (the old decision-model input) 0.000 W 9.6 W 0.06 W
MiniLM Core ML encoder (new) 1.84 W 3.3 W 0.06 W
Vision saliency + classify 0.69 W 5.1 W 0.11 W
Apple Intelligence decide 4.67 W 1.8 W 0.07 W

Apple's contextual embedding runs on the CPU, not the Neural Engine. The decision model now uses all-MiniLM-L6-v2, converted to an fp16 Core ML program: 0.76 ms per sentence, against 8 ms for the old embedding. symbio_ane/build_text_encoder.py builds it into <home>/cache/text-encoder from a separate coremltools venv. Its output matches the torch model at cosine 0.99997.

Decision order, with think/no-think accuracy on 31 held-out messages:

source runs on score latency
MiniLM vote ANE 30/31 <1 ms
Apple Intelligence's on-device model ANE 27/31 ~0.7 s
CPU embedding vote CPU 30/31 ~8 ms
regex — 19/31 —

In live turns on the 14B, the decision came from minilm-ane in 2 ms.

huyedits and others added 3 commits September 26, 2026 23:30
…edding was on the cpu T_T so the decision model is now minilm as a core ml program on the neural engine (<1ms, 1.8W ane, 30/31 held out). apple intelligence is the next opinion (27/31, 0.7s, also ane) >//<

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ryuzz9EtCjeeR8ofVxhGW
…st-open-a-page so the turn stopped after x.com loaded T_T now only a bare destination ends it, "continue" keeps the task's thinking, and the window keeps following the chat past 5 msgs instead of hiding each answer under the fold :3

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ryuzz9EtCjeeR8ofVxhGW
@huyedits
huyedits merged commit ed8e436 into main Sep 26, 2026
0 of 2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant