Tools · Audio Tools

Video & Audio Text Extractor

Pull spoken words, on-screen text, or both from a file on this device. Decoding depends on this browser.

Drop audio or video

Audio: MP3, WAV, M4A, AAC, OGG, WebM. Video: MP4, WebM, MOV, M4V if this browser can decode it.

Size — · Duration —

Uses the existing local Whisper-compatible ONNX model for speech and Tesseract.js for on-screen frames. No OpenAI, Azure, Google, or Amazon transcription or OCR API. ffmpeg.wasm is not used. API: None.

Read the original guide: Extract speech and on-screen text without a cloud API.

An unhandled error has occurred. Reload Dismiss

Rejoining the server...

Rejoin failed... trying again in seconds.

Failed to rejoin.
Please retry or reload the page.

The session has been paused by the server.

Failed to resume the session.
Please retry or reload the page.