Audio to Text
This audio to text tool never uploads your recording. It decodes the file in this browser tab, and only when you press the button that names the size does it download 71.24 MB of speech model and runtime (Whisper tiny.en, English only) from this site, check every file against a pinned SHA-256 and run the model in a background worker on your device. Files up to 50 MiB and 5 minutes are transcribed in 30-second windows with a 5-second overlap, one window at a time, and the result is a draft you edit: every segment has its own start and end, a Play button that replays that part of the original audio, and text you can correct. An end the model did not give stays unset until you set it, SRT and VTT export wait until every time is valid, TXT and JSON never wait on the times, and a draft saved here as JSON, SRT or VTT opens again to carry on editing.
Open or drop a short English recording, or a video whose sound your browser can decode. It is decoded on this device and never uploaded.
model Whisper tiny.en, English only
download 71.24 MB, only when you press the button,
every file checked before useThe first press downloads 71,240,073 bytes of model and runtime from this site, checks each file against its pinned SHA-256, then transcribes here. Your browser may reuse a copy from an earlier visit.
The draft appears here as timed segments you can play, correct and export. A draft saved here as .json, .srt or .vtt opens again to carry on editing.
- Clip
- none
- Segments
- 0
- Ends unset
- 0
- Export
- TXT style
There is no text to put in subtitles yet.
- a small blue boat waits beside the quiet river. Please bring three notebooks to the meeting. One segment, 0:00.00 to 0:05.24recorded in testing
- no detectable signal, the model is not called (run raw, it wrote "you")
- flagged, end before start; SRT and VTT wait until it is fixed
Model, runtime and licences
Speech recognition is Whisper tiny.en by OpenAI (MIT licence), in the ONNX conversion Xenova/whisper-tiny.en at revision 79fb389 (16 December 2025, Apache-2.0 declaration), run by Transformers.js 4.3.0 (Apache-2.0) on ONNX Runtime Web 1.31.0-dev.20260914-8d85527a0 (MIT), single-threaded WebAssembly. Every file is served from this site; the files a transcription needs are checked against the SHA-256 values below before use. Files are limited to 52.43 MB (50 MiB) and 5 minutes.
| File | Bytes | Licence | SHA-256 |
|---|---|---|---|
| config.json | 2,202 | MIT original weights; Apache-2.0 ONNX conversion declaration | 37a1073be00d19118c06557896c7c148598f4d8277edc0f5bc07c9f5554839f1 |
| tokenizer_config.json | 835 | MIT original weights; Apache-2.0 ONNX conversion declaration | e082c1ad251541bf277967a703252cddd4bb37a71a43737e03d050c22ec08238 |
| preprocessor_config.json | 339 | MIT original weights; Apache-2.0 ONNX conversion declaration | a6a76d28c93edb273669eb9e0b0636a2bddbb1272c3261e47b7ca6dfdbac1b8d |
| tokenizer.json | 2,128,494 | MIT original weights; Apache-2.0 ONNX conversion declaration | c6ee8f089220a5b1188f6426456772572671c6141ae007eecb83c6a8349f5deb |
| generation_config.json | 1,590 | MIT original weights; Apache-2.0 ONNX conversion declaration | 132c95ba9db45f4498f2eab3fea7c1d6a174005010f8f6b7d20cfd5e9795996b |
| encoder_model_quantized.onnx | 10,124,913 | MIT original weights; Apache-2.0 ONNX conversion declaration | 8cc3c6f8563d1b3fbd2c5af9f64c2bed8b020bc593c402d1ef53b9f08fbf1b90 |
| decoder_model_merged_quantized.onnx | 30,727,382 | MIT original weights; Apache-2.0 ONNX conversion declaration | dbb2e063b7fbc41d9803b9698f93ecb035c50cbb3fb87b56cb131e4a5eb99059 |
| README.md | 3,213 | MIT original weights; Apache-2.0 ONNX conversion declaration | 2f4ddb95ec9d6c6d5f6a1374fc004f45e91f6d7f450ee479f765b8aca47e8b4d |
| huggingface-transformers-transformers.js | 1,339,484 | Apache-2.0 | 43f86bbcab7cb088c6bd740e2bdec12ae9c09cc1c63329e58769f5a1e7c70bf7 |
| huggingface-transformers-LICENSE | 11,358 | Apache-2.0 | cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 |
| onnxruntime-web-ort-wasm-simd-threaded.asyncify.mjs | 53,057 | MIT | 0966b6105cd936744498aa60df7a22cbd47af3374dbc64a9ab561c08a71e3611 |
| onnxruntime-web-ort-wasm-simd-threaded.asyncify.wasm | 26,861,777 | MIT | 49871f5a4409519797e127440868a6d1923339d9185907f301a5b2a1d90af082 |
| whisper-README.md | 8,246 | MIT license evidence | 38c180c2a8d8ba628a3e131ad42bf06c8a51bb681efb30fbe28d5a0f4b65f288 |
| whisper-LICENSE | 1,063 | MIT | b5d65a59060e68c4ff940e1eddfa6f94b2d68fdf58ed7f4dd57721c997e35e9d |
| onnxruntime-LICENSE | 1,073 | MIT | 2f07c72751aed99790b8a4869cf2311df85a860b22ded05fa22803587a48922c |
| onnxruntime-ThirdPartyNotices.txt | 338,088 | Retain complete bundled third-party notices | 143764b952fdb1a7c69ce653bfba74a7744d6a8a573bfb73e235fba356c83de3 |
| dependency-notices.json | 52,788 | Apache-2.0 / MIT / BSD-3-Clause / ISC dependency declarations and retained notices | 5a3598d4883e45b1b99a15ad75a05cfccd33b30c210ccf1645b8afb849fcdd45 |
| NOTICE.txt | 558 | Original attribution notice | 57894bd08cbf1f4eff63460e2786b59eebb626114512f6b51fccf83b6263915a |
Common questions
- Is my recording uploaded anywhere?
- No. The file is decoded and transcribed in this browser tab, and neither the audio nor the draft is sent anywhere. The only download is the speech model and runtime, 71.24 MB from this site, fetched when you press the button that names the size. The page remembers one setting on this device, the TXT style, and never your file or your text.
- How accurate is the transcript?
- It is a draft from Whisper tiny.en, the smallest English-only Whisper model, and there is no accuracy score. It can mishear words, and it can invent words over noise or silence: given five seconds of digital silence directly, it wrote "you". That is why audio whose samples are all exactly zero is skipped without calling the model, and why every segment has a Play button: listen, then correct the text before you use it. Timestamps are approximate, and a segment that follows a silent stretch can be given a start inside the silence.
- Which files can I open?
- An audio file your browser can decode, such as MP3, WAV or M4A, or a video file whose sound your browser can decode, such as many MP4 and WebM files. The limit is 50 MiB and 5 minutes; for a longer recording, trim it or cut it into parts with an audio splitter first. The speech must be English: this model has no other language.
- Why is the download 71.24 MB, and does it happen every time?
- The ten files are the Whisper tiny.en encoder and decoder (40.85 MB together), its tokenizer and settings, the Transformers.js 4.3.0 runtime and ONNX Runtime's WebAssembly. They are fetched from this site only when you press the button, each checked against its pinned SHA-256 before use, and the model stays loaded until you cancel, a window fails or you leave the page, so a second file does not download it again. On a later visit your browser may reuse a copy it kept, and every file is checked again.
- Can I make SRT or VTT subtitles?
- Yes, once every segment has a valid start and end. A segment end the model did not give is left unset rather than guessed, so set it first. An end before its start, a time past the clip's end or a start before the previous segment's start is flagged, and the SRT and VTT buttons wait until each is fixed. TXT, plain or with times, and JSON do not wait. These are plain subtitle files made from a draft, not certified captions.
- How are recordings longer than 30 seconds handled?
- The audio is cut into 30-second windows with a 5-second overlap and transcribed one window at a time, and a window still running after 60 seconds is stopped. Where two windows overlap, a piece of text is removed only when all of its words also appear in the other window's text for that overlap, so words only one window heard are kept, and anything left that repeats is flagged for you to check. A window that is exactly silent is skipped. If a window fails or you cancel, the finished windows stay as a draft marked partial.
- Does it tell speakers apart or translate?
- No. There is no speaker labelling, diarization or translation: the model writes English text from English speech, and the draft is one list of timed segments.
Whisper tiny.en, a small English-only speech model, writes a draft in your browser that can mishear words and can invent them over noise or silence; each segment has a Play button for checking it against the recording. Audio whose samples are all exactly zero is skipped without calling the model, and an end time the model did not give stays unset until you set it, never guessed. Timestamps are approximate, and there is no accuracy score, speaker labelling, translation or certified captioning.