Articles ยท Audio Tools

Extract speech and on-screen text without a cloud API

Lectures, product demos, and recorded stand-ups hide two kinds of text: what people said and what was on the slide. Students need a study transcript. Teachers need captions. Support teams need a record of a call without pasting the file into a cloud speech product. A slide-heavy walkthrough also needs the heading that never got spoken. DevNestro’s Video & Audio Text Extractor keeps the file on this device and reuses the speech and OCR engines already on the site instead of shipping a second Whisper or a second Tesseract.

Why this is worth doing in the browser

Hosted speech and OCR APIs are convenient until the file is a hiring interview, a medical discussion, or an internal deck. They also require keys, billing, and a network path that copies the media. This page decodes supported formats in the browser, lazy-loads the existing Whisper-compatible ONNX model only for Spoken words or Both, and lazy-loads Tesseract.js only for Visible text or Both. WebGPU is preferred for speech, with a WASM fallback. It does not call OpenAI, Azure Speech, Google Speech, AWS Transcribe, or a cloud OCR API. It does not load ffmpeg.wasm, does not split speakers, and does not claim a complete transcript. Not every codec works in every browser; the page tells you when the file cannot be read locally. Audio-only files hide on-screen OCR. Video can run Spoken, Visible, or Both. Frames are sampled on an interval with a similar-frame skip in Smart Scan—never every frame.

How to use the tool

Drop an MP3, WAV, M4A, AAC, OGG, or WebM audio file, or an MP4, WebM, MOV, or M4V video if this browser can decode it. Read size and duration in the player, optionally set a time range, then choose Spoken words, Visible text, or Both. For on-screen text, pick Smart Scan, Subtitles (bottom 25%), Screen/presentation, or a custom region with presets (full, top half, center, lower third, bottom 25%). Extract. Watch real progress: a percent only when the speech model reports one, otherwise stages such as preparing audio, and for OCR “frame n of m.” Cancel stops the current run; loaded models may stay in the tab. Edit the text without forcing a new run. Combined Clean Text drops near-duplicate lines that appear in both speech and OCR; the Spoken and Visible tabs keep the original evidence. Copy or download TXT, timestamped TXT, CSV, SRT, or VTT. Search All, Spoken, or On screen and click a timeline cue to seek. Open Text to Speech with spoken, visible, or combined text. First use can take longer while a model or language pack downloads.

  1. Choose a file the current browser can decode; do not assume every MP4 or MOV codec works everywhere.
  2. Load speech or OCR only after you click Extract so other pages stay light.
  3. Use Smart Scan or a subtitle crop instead of expecting every frame to be read.
  4. Proofread before you treat the output as minutes or captions of record; accuracy is not 100%.

Privacy

Media is processed on this device and is not uploaded to a transcription or OCR service. Speech models and OCR language packs may still be downloaded from a trusted host on first use. Filenames and extracted text are not sent to analytics.

Need captions, slide text, or minutes from a file you would not upload? Open Video & Audio Text Extractor, extract locally, and download SRT or TXT from the same page.

Open Video & Audio Text Extractor

An unhandled error has occurred. Reload Dismiss

Rejoining the server...

Rejoin failed... trying again in seconds.

Failed to rejoin.
Please retry or reload the page.

The session has been paused by the server.

Failed to resume the session.
Please retry or reload the page.