Related to #879
#879 / #880 fix the case where microphone setup rejects during streamFromMicrophone(). This issue is about a related but distinct case: microphone setup can hang forever (neither resolving nor rejecting) with no timeout anywhere in the chain. Even after #880 lands, this case would still produce a silently stuck connection, since there's no rejection for the new error-forwarding to catch.
Where
packages/client/src/platform/web/webScribeMicrophoneSetup (bundled as webScribeMicrophoneSetup in dist/lib.iife.js):
const source = audioContext.createMediaStreamSource(stream);
const scribeNode = new AudioWorkletNode(audioContext, "scribeAudioProcessor");
...
source.connect(scribeNode);
if (audioContext.state === "suspended") await audioContext.resume();
This await audioContext.resume() has no timeout and no fallback. It's called from ScribeRealtime.streamFromMicrophone(), itself fired fire-and-forget off the WebSocket "open" event (scribe.ts), fully detached from the Scribe.connect() return value and from RealtimeConnection's event lifecycle.
The problem
In some browsers (most reproducibly Safari/WebKit), AudioContext.resume() can fail to settle at all when called too many awaits away from the originating user gesture — it neither resolves nor rejects, it just never returns. Since nothing in this call chain has a timeout:
- streamFromMicrophone() hangs forever on this line.
- No RealtimeEvents.ERROR fires (nothing rejected).
- The WebSocket itself can still be open and can still receive session_started from the server (that's independent of whether the client ever streams audio back), so useScribe()'s status can show "connected" and onSessionStarted can fire normally.
- From the consumer's perspective, the connection looks completely healthy — no error, no disconnect — while zero audio is ever captured or sent. There is no event or state change a caller can observe to detect this.
Expected behavior
Wrap the microphone setup (or at minimum the audioContext.resume() call) in a bounded timeout. On timeout, treat it the same as a setup failure — reject with a clear error (e.g. "Microphone setup timed out
waiting for AudioContext to resume") and route it through the same error-reporting path as ROR fires and callers have something to react to instead of an unrecoverable, undetectablehang.
Workaround on our side
We're currently working around this at the application layer with a ~7s "connect confirmed" watchdog that requires onSessionStarted and actual transcript activity within a bounded window after connect() resolves, otherwise we tear down and retry since there's no SDK-level signal to rely on. let us remove that. ;)
Greetings and Thanks for your help in advance!
Related to #879
#879 / #880 fix the case where microphone setup rejects during
streamFromMicrophone(). This issue is about a related but distinct case: microphone setup can hang forever (neither resolving nor rejecting) with no timeout anywhere in the chain. Even after #880 lands, this case would still produce a silently stuck connection, since there's no rejection for the new error-forwarding to catch.Where
packages/client/src/platform/web/webScribeMicrophoneSetup(bundled aswebScribeMicrophoneSetupindist/lib.iife.js):This await audioContext.resume() has no timeout and no fallback. It's called from ScribeRealtime.streamFromMicrophone(), itself fired fire-and-forget off the WebSocket "open" event (scribe.ts), fully detached from the Scribe.connect() return value and from RealtimeConnection's event lifecycle.
The problem
In some browsers (most reproducibly Safari/WebKit), AudioContext.resume() can fail to settle at all when called too many awaits away from the originating user gesture — it neither resolves nor rejects, it just never returns. Since nothing in this call chain has a timeout:
Expected behavior
Wrap the microphone setup (or at minimum the audioContext.resume() call) in a bounded timeout. On timeout, treat it the same as a setup failure — reject with a clear error (e.g. "Microphone setup timed out
waiting for AudioContext to resume") and route it through the same error-reporting path as ROR fires and callers have something to react to instead of an unrecoverable, undetectablehang.
Workaround on our side
We're currently working around this at the application layer with a ~7s "connect confirmed" watchdog that requires onSessionStarted and actual transcript activity within a bounded window after connect() resolves, otherwise we tear down and retry since there's no SDK-level signal to rely on. let us remove that. ;)
Greetings and Thanks for your help in advance!