Skip to content

Qwen3-ASR mishandles invalid, short, and no-speech audio inputs #928

Description

@ainergiz

Problem

Qwen3-ASR preprocessing loses the distinction between invalid or short input and mlx-audio's default min_chunk_duration=1.0 padding:

  • A zero-frame array is padded and sent to the model, which can hallucinate text.
  • NaN and infinite samples are not rejected before preprocessing.
  • A valid 10 ms clip is reported with a 1.0-second segment end after padding.
  • An empty language None<asr_text> response exposes the literal language None instead of an empty no-speech result.
  • Auto-detected language from one chunk is reused as the request for later chunks, so a silent chunk can corrupt following transcription.

Expected behavior

  • Reject empty and non-finite Qwen3-ASR audio before inference.
  • Keep minimum-length padding for valid short clips, but report their original duration.
  • Parse language None<asr_text> as empty language and text.
  • Keep automatic language detection independent for every chunk.

A real-model check reproduced the empty-input hallucination and padded duration. Tested on current main at 246e233.

Activity

  1. linear-code commented on Aug 29, 2026

    @linear-code
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions