Drop a recording of up to 30 minutes. It is transcribed here, in this tab — nothing is uploaded and there is no account. Expect roughly a quarter of the recording's length in waiting; you can watch the text appear as it goes.
or click to choose one · mp3, m4a, wav, mp4, mov and more
Nothing is uploaded — it is transcribed on this device.
No. The speech model is downloaded to your browser and runs there. Your file is read by the page and never sent to a server — you can disconnect from the internet after the model loads and it still works.
The first file waits for an ~80 MB model download. Your browser caches it, so every run after that starts immediately. Machines with WebGPU transcribe several times faster than those falling back to WASM.
Up to 30 minutes. It is processed in 30-second windows at roughly four times real time, so a 20-minute recording takes about five minutes — you can watch the transcript appear while it works, and the tab can sit in the background. Longer than that and the wait stops being reasonable, which is what the app is for.
Good for clear speech, imperfect on heavy accents, crosstalk, and names it has never seen. It also does not separate speakers. If you need that, the Kika app uses a much larger model and labels who said what.
This page sees one file, in isolation, and forgets it when you close the tab. Kika sits on your side of every call — no bot joins — and afterwards connects what was said to everything said before it.
Five minutes before your next call, that becomes a brief — every line attributed to the meeting and date it came from, so you can check it rather than trust it. How we build it →