gizmobench

Audio to Text

This audio to text tool never uploads your recording. It decodes the file in this browser tab, and only when you press the button that names the size does it download 71.24 MB of speech model and runtime (Whisper tiny.en, English only) from this site, check every file against a pinned SHA-256 and run the model in a background worker on your device. Files up to 50 MiB and 5 minutes are transcribed in 30-second windows with a 5-second overlap, one window at a time, and the result is a draft you edit: every segment has its own start and end, a Play button that replays that part of the original audio, and text you can correct. An end the model did not give stays unset until you set it, SRT and VTT export wait until every time is valid, TXT and JSON never wait on the times, and a draft saved here as JSON, SRT or VTT opens again to carry on editing.

audio or video fileup to 50 MiB and 5 minutes

Open or drop a short English recording, or a video whose sound your browser can decode. It is decoded on this device and never uploaded.

model     Whisper tiny.en, English only
download  71.24 MB, only when you press the button,
          every file checked before use

The first press downloads 71,240,073 bytes of model and runtime from this site, checks each file against its pinned SHA-256, then transcribes here. Your browser may reuse a copy from an earlier visit.

draftempty

The draft appears here as timed segments you can play, correct and export. A draft saved here as .json, .srt or .vtt opens again to carry on editing.

Clip
none
Segments
0
Ends unset
0
Export
TXT style

There is no text to put in subtitles yet.

  • Synthetic English test clip, 5.498 s16 words, two sentences
    a small blue boat waits beside the quiet river. Please bring three notebooks to the meeting. One segment, 0:00.00 to 0:05.24recorded in testing
  • 5.00 s of digital silenceevery sample exactly 0
    no detectable signal, the model is not called (run raw, it wrote "you")
  • End typed as 0:04.00on a segment starting 0:04.62
    flagged, end before start; SRT and VTT wait until it is fixed
Model, runtime and licences

Speech recognition is Whisper tiny.en by OpenAI (MIT licence), in the ONNX conversion Xenova/whisper-tiny.en at revision 79fb389 (16 December 2025, Apache-2.0 declaration), run by Transformers.js 4.3.0 (Apache-2.0) on ONNX Runtime Web 1.31.0-dev.20260914-8d85527a0 (MIT), single-threaded WebAssembly. Every file is served from this site; the files a transcription needs are checked against the SHA-256 values below before use. Files are limited to 52.43 MB (50 MiB) and 5 minutes.

FileBytesLicenceSHA-256
config.json2,202MIT original weights; Apache-2.0 ONNX conversion declaration37a1073be00d19118c06557896c7c148598f4d8277edc0f5bc07c9f5554839f1
tokenizer_config.json835MIT original weights; Apache-2.0 ONNX conversion declaratione082c1ad251541bf277967a703252cddd4bb37a71a43737e03d050c22ec08238
preprocessor_config.json339MIT original weights; Apache-2.0 ONNX conversion declarationa6a76d28c93edb273669eb9e0b0636a2bddbb1272c3261e47b7ca6dfdbac1b8d
tokenizer.json2,128,494MIT original weights; Apache-2.0 ONNX conversion declarationc6ee8f089220a5b1188f6426456772572671c6141ae007eecb83c6a8349f5deb
generation_config.json1,590MIT original weights; Apache-2.0 ONNX conversion declaration132c95ba9db45f4498f2eab3fea7c1d6a174005010f8f6b7d20cfd5e9795996b
encoder_model_quantized.onnx10,124,913MIT original weights; Apache-2.0 ONNX conversion declaration8cc3c6f8563d1b3fbd2c5af9f64c2bed8b020bc593c402d1ef53b9f08fbf1b90
decoder_model_merged_quantized.onnx30,727,382MIT original weights; Apache-2.0 ONNX conversion declarationdbb2e063b7fbc41d9803b9698f93ecb035c50cbb3fb87b56cb131e4a5eb99059
README.md3,213MIT original weights; Apache-2.0 ONNX conversion declaration2f4ddb95ec9d6c6d5f6a1374fc004f45e91f6d7f450ee479f765b8aca47e8b4d
huggingface-transformers-transformers.js1,339,484Apache-2.043f86bbcab7cb088c6bd740e2bdec12ae9c09cc1c63329e58769f5a1e7c70bf7
huggingface-transformers-LICENSE11,358Apache-2.0cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30
onnxruntime-web-ort-wasm-simd-threaded.asyncify.mjs53,057MIT0966b6105cd936744498aa60df7a22cbd47af3374dbc64a9ab561c08a71e3611
onnxruntime-web-ort-wasm-simd-threaded.asyncify.wasm26,861,777MIT49871f5a4409519797e127440868a6d1923339d9185907f301a5b2a1d90af082
whisper-README.md8,246MIT license evidence38c180c2a8d8ba628a3e131ad42bf06c8a51bb681efb30fbe28d5a0f4b65f288
whisper-LICENSE1,063MITb5d65a59060e68c4ff940e1eddfa6f94b2d68fdf58ed7f4dd57721c997e35e9d
onnxruntime-LICENSE1,073MIT2f07c72751aed99790b8a4869cf2311df85a860b22ded05fa22803587a48922c
onnxruntime-ThirdPartyNotices.txt338,088Retain complete bundled third-party notices143764b952fdb1a7c69ce653bfba74a7744d6a8a573bfb73e235fba356c83de3
dependency-notices.json52,788Apache-2.0 / MIT / BSD-3-Clause / ISC dependency declarations and retained notices5a3598d4883e45b1b99a15ad75a05cfccd33b30c210ccf1645b8afb849fcdd45
NOTICE.txt558Original attribution notice57894bd08cbf1f4eff63460e2786b59eebb626114512f6b51fccf83b6263915a

Common questions

Is my recording uploaded anywhere?
No. The file is decoded and transcribed in this browser tab, and neither the audio nor the draft is sent anywhere. The only download is the speech model and runtime, 71.24 MB from this site, fetched when you press the button that names the size. The page remembers one setting on this device, the TXT style, and never your file or your text.
How accurate is the transcript?
It is a draft from Whisper tiny.en, the smallest English-only Whisper model, and there is no accuracy score. It can mishear words, and it can invent words over noise or silence: given five seconds of digital silence directly, it wrote "you". That is why audio whose samples are all exactly zero is skipped without calling the model, and why every segment has a Play button: listen, then correct the text before you use it. Timestamps are approximate, and a segment that follows a silent stretch can be given a start inside the silence.
Which files can I open?
An audio file your browser can decode, such as MP3, WAV or M4A, or a video file whose sound your browser can decode, such as many MP4 and WebM files. The limit is 50 MiB and 5 minutes; for a longer recording, trim it or cut it into parts with an audio splitter first. The speech must be English: this model has no other language.
Why is the download 71.24 MB, and does it happen every time?
The ten files are the Whisper tiny.en encoder and decoder (40.85 MB together), its tokenizer and settings, the Transformers.js 4.3.0 runtime and ONNX Runtime's WebAssembly. They are fetched from this site only when you press the button, each checked against its pinned SHA-256 before use, and the model stays loaded until you cancel, a window fails or you leave the page, so a second file does not download it again. On a later visit your browser may reuse a copy it kept, and every file is checked again.
Can I make SRT or VTT subtitles?
Yes, once every segment has a valid start and end. A segment end the model did not give is left unset rather than guessed, so set it first. An end before its start, a time past the clip's end or a start before the previous segment's start is flagged, and the SRT and VTT buttons wait until each is fixed. TXT, plain or with times, and JSON do not wait. These are plain subtitle files made from a draft, not certified captions.
How are recordings longer than 30 seconds handled?
The audio is cut into 30-second windows with a 5-second overlap and transcribed one window at a time, and a window still running after 60 seconds is stopped. Where two windows overlap, a piece of text is removed only when all of its words also appear in the other window's text for that overlap, so words only one window heard are kept, and anything left that repeats is flagged for you to check. A window that is exactly silent is skipped. If a window fails or you cancel, the finished windows stay as a draft marked partial.
Does it tell speakers apart or translate?
No. There is no speaker labelling, diarization or translation: the model writes English text from English speech, and the draft is one list of timed segments.

Whisper tiny.en, a small English-only speech model, writes a draft in your browser that can mishear words and can invent them over noise or silence; each segment has a Play button for checking it against the recording. Audio whose samples are all exactly zero is skipped without calling the model, and an end time the model did not give stays unset until you set it, never guessed. Timestamps are approximate, and there is no accuracy score, speaker labelling, translation or certified captioning.