100% on-device inference

Speech to text, running in your browser

OpenAI's Whisper, executed locally with WebGPU and ONNX Runtime Web. Your audio is decoded in the tab and never uploaded — there is no backend to upload it to.

No server. No API key. 99+ languages WebGPU or WASM fallback Export .txt / .srt

Input

Drop an audio file or click to browse · wav, mp3, m4a, ogg, flac, webm
Or try:
Recording 0:00 — click Record again to stop
idle
— 0.0 s · 0 Hz · mono

Engine log

device: detecting…
ready.

Transcript

Audio
—
Compute
—
Speed
—
Words
—

Segments — click to seek

How this works

Nothing leaves the tab. The file is read with the Web Audio API, mixed to mono and resampled to 16 kHz with an OfflineAudioContext, then fed straight into ONNX Runtime Web. There is no fetch of your audio and no endpoint behind this page.

Precision. On WebGPU the encoder runs in fp32 and the decoder in q4; on the WASM path both run q8. Weights are fetched from the Hub on first use and cached by the browser afterwards — watch the log for per-file progress.

Built on transformers.js v4 · onnx-community/whisper-base · OpenAI Whisper