Pull spoken words, on-screen text, or both from a file on this device. Decoding depends on this browser.
Drop audio or video
Audio: MP3, WAV, M4A, AAC, OGG, WebM. Video: MP4, WebM, MOV, M4V if this browser can decode it.
Size — · Duration —
Player
Click a timeline cue after extraction to seek. For a custom OCR region, choose Custom and drag on the video.
What to extract
Audio-only files use speech recognition. Visible text and Both appear for video.
Region presets
Frames are sampled, not every frame. Similar successive OCR lines are merged. This is not 100% capture.
The first use of speech or OCR may take longer while the model or language pack downloads. Whisper loads only for Spoken or Both. Tesseract loads only for Visible text or Both.
Preparing…
Combined Clean Text drops near-duplicate lines that appear in both speech and on-screen text. Spoken and Visible tabs keep the original evidence.
Timeline
No cues yet.
Uses the existing local Whisper-compatible ONNX model for speech and Tesseract.js for on-screen frames. No OpenAI, Azure, Google, or Amazon transcription or OCR API. ffmpeg.wasm is not used. API: None.