Skip to content

ASR: migrate SenseVoiceSmall to int8 sherpa-onnx runtime (drop PyTorch/funasr) #135

Description

@BrettKinny

Motivation

Follow-up from #124. The no-GPU ASR default runs SenseVoiceSmall via the funasr
AutoModel path, which pulls in PyTorch + torchaudio and a ~900 MB model.pt.
That stack is heavy and fragile on small/ARM hosts (the #124 crash was a tokenizer
asset that funasr's loader resolved to bpemodel=None), and we have no measured
CPU/ARM real-time factor
for the eager funasr path.

The same model served via int8 sherpa-onnx / ONNX Runtime is dramatically
lighter and faster on exactly the hosts that hit #124:

  • ~230 MB int8 ONNX weights (vs ~900 MB model.pt)
  • no PyTorch dependency — onnxruntime only, with first-class aarch64 wheels
  • ~10–20× real-time on ARM Cortex-A76 (sherpa-onnx RK3588 CPU test measured
    SenseVoice int8 at RTF ~0.10→0.05 @ 1→4 threads; the Pi 5 is 4× A76 @2.4 GHz,
    same core family) — i.e. comfortably real-time on a Pi 5 without a GPU
  • removes the funasr version-pin + the fun_local.py monkey-patch + the bpemodel
    asset fragility that caused xiaozhi-esp32-server crash looping #124

Scope

  • Add a sherpa-onnx-based ASR provider (or adapt fun_local.py) that loads the
    int8 ONNX SenseVoiceSmall export.
  • Update make fetch-models to pull the ONNX assets; update make doctor checks.
  • Keep language: en behaviour (the existing patch's purpose).
  • Benchmark on a real Pi 5 / CPU host and record the RTF before flipping the default.

Notes

Refs #124, #47.

Activity

  1. added
    enhancementNew feature or request
    area:xiaozhiUnraid xiaozhi-server container + custom providers
    on Jun 3, 2026
  2. BrettKinny commented on Jun 3, 2026

    @BrettKinny
    OwnerAuthor

    Opt-in provider landed in #140 (SenseVoiceOnnx, type sensevoice_onnx) — coexists with FunASR, which stays the no-GPU default for now.

    Done:

    • sherpa-onnx int8 provider mirroring whisper_local.py; language: en passed natively (no lang_tag_filter / English-pin patch).
    • make fetch-models + make doctor + dotty_doctor.py handle models/SenseVoiceSmall-onnx/ (model.int8.onnx + tokens.txt).
    • sherpa-onnx==1.13.2 in the Dockerfile (onnxruntime bundled — no torch).
    • Verified: clean import + accurate transcript in an ephemeral container; both compose files validate.

    Provisional RTF (Docker host = Intel i5-3570, 4-core, 2012-era — weaker than the Pi-5 target, so a conservative floor), 7.15 s English utterance:

    num_threads RTF speedup
    1 0.226 ~4.4×
    2 (default) 0.123 ~8.2×

    Comfortably real-time without a GPU.

    Still open (gates the default-flip):

    • On-target Pi 5 / RK3588-A76 RTF (the ~0.05–0.10 figure this issue targets). The provider logs an ASR-RTF provider=sensevoice_onnx … line per utterance, so this can be read straight from production logs once deployed. (Also asked the original xiaozhi-esp32-server crash looping #124 reporter for an ARM number.)
    • Flip the no-GPU default FunASR → SenseVoiceOnnx (separate PR, citing the on-target number).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:xiaozhiUnraid xiaozhi-server container + custom providersenhancementNew feature or requeststatus:speculativeSpeculative / future capability

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions