Transcribe audio to text

Use your browser's speech recognition to convert audio or video into SRT, VTT, or plain text. Real-time, private, free.

How to Use

  1. Choose the spoken language of your audio.
  2. Select an audio or video file (up to 100MB).
  3. The audio plays once while the browser transcribes it in real time.
  4. Download the transcript as SRT, VTT, or plain text.
  5. Verify the output file in your device's download folder.
  6. You can adjust settings and process new files immediately without any restrictions.

All processing is done in your browser, and files are never sent to a server.

Frequently Asked Questions

In your browser, through the built-in Web Speech API. Some implementations may use cloud recognition (browser-dependent). Files are not uploaded by ConvertBox.
The Web Speech API listens to your device's audio output as the file plays. Mute the speakers if you prefer β€” the recognition still works.
Accuracy is best on clear single-speaker recordings. Background noise, accents, and overlapping speech reduce quality.
No. All operations are processed entirely inside your browser's local sandbox using client-side APIs (like WebAssembly, pdf-lib, or HTML5 Canvas). Your files and data never leave your device.
The tool is completely free with no usage limits. However, since processing occurs in your browser's memory, we recommend keeping file sizes under 50MB (especially for heavy files like PDFs or videos) to ensure stability.
No. All tools on ConvertBox are 100% free. We do not insert watermarks, limit features, or require any registration or payment.
Yes. Once the page and its local library dependencies are loaded in your browser, the conversion logic runs completely offline without needing an active internet connection.