Free tool · No sign-up · Nothing uploaded

Audio to text, in your browser

Drop in a recording and get a transcript with timestamps. It runs on your own device with an open-source speech model, so there’s nothing to upload, nobody to sign up with, and no minutes to count.

Speech model · startingWhisper base

Drop an audio or video file

or browse · MP3, WAV, M4A, OGG, FLAC, MP4, MOV

Whisper base · good for most recordings · ~130 MB, downloaded once

Works in current Chrome, Edge, Safari, and Firefox. The speech model downloads once (from ~40 MB) and stays cached.

How it works

A transcript in three moves

The whole thing happens on this page. There’s no queue, no email with a link, and no server on our end that ever hears your recording.

01

Drop in a recording

A voice memo, a lecture, an interview, a meeting, the audio track of a video. Or hit Record and talk. Any format your browser can play.

02

Your browser does the work

An open-source speech model downloads once and runs on your own device — on your GPU where the browser has one. The audio never leaves your machine.

03

Take the transcript with you

Read it as text or as timestamped lines you can click to replay. Copy it, or download TXT, SRT, or VTT subtitles.

Why it’s free

Free because it costs us nothing to run

Paid transcription tools charge by the minute because they pay for the servers your audio runs on. This one runs on yours — so the price can be honest.

Nothing leaves your device

The transcription runs inside your browser. We never receive the audio or the text — there's nothing for us to store, sell, or leak.

No account, no clock

No sign-up, no trial minutes, no watermark, no “upgrade to export”. Every result is yours to copy or download, every time.

Speaks 99 languages

OpenAI's Whisper, the model most paid tools are built on, detects the language automatically — or set it, and pick the model size that suits the recording.

FAQ

Questions, answered

Is it really free? What's the catch?+

It's free, with no account, no trial minutes, and no watermark. There's no catch because there's no cost on our side: the transcription runs in your browser on your own device, using an open-source speech model, so we're not paying for servers to process your audio. We built it because people who need a transcript today are often the same people who'd like the transcript written while the call is still happening — which is what Assistly, our desktop app, does. The tool is the free version of the idea.

Where does my audio go?+

Nowhere. The file is decoded and transcribed inside your browser — nothing is uploaded, and we never see the audio or the transcript. The only network request is the one-time download of the speech model from the Hugging Face Hub, which your browser then caches. You can confirm this yourself: start a transcription, switch off Wi-Fi once the model has loaded, and it keeps going.

What files can I transcribe?+

Anything your browser can play: MP3, WAV, M4A/AAC, OGG, FLAC, Opus, and WebM audio, plus the soundtrack of MP4 and MOV video. You can also record straight from your microphone. There's no file-size or length limit imposed by the tool; very long recordings are limited by your device's memory, so for anything over a couple of hours split the file first.

How accurate is it, and what model is it using?+

It runs OpenAI's Whisper, the open-source speech model most transcription products are built on, in three sizes. Fast (Whisper tiny) is quick and fine for clean speech; Balanced (Whisper base) is the default and handles most recordings well; Accurate (Whisper small) makes the fewest mistakes, especially on accents, names, and background noise, at the cost of a longer wait. All three recognise 99 languages and detect the language automatically, or you can set it. Expect the results to be roughly on par with the free tiers of paid services, and to improve if you pick a bigger model.

Why does the first run take a while?+

The speech model has to be downloaded once — about 40 MB for Fast, 80 MB for Balanced, and 250 MB for Accurate. Your browser keeps it cached, so every transcription after that starts immediately. On a computer with a modern GPU the transcription itself typically runs several times faster than real time; on older hardware or phones it can take about as long as the recording.

Can it tell speakers apart?+

Not yet — Whisper transcribes what was said, not who said it, so the transcript is one voice with timestamps. If you need speaker labels for meetings, that's one of the things the Assistly desktop app does: it separates your voice from the other participants and tracks who said what, live, then writes the recap and action items when the call ends.

Which browsers does it work in?+

Current versions of Chrome, Edge, Safari, and Firefox on desktop, and most modern mobile browsers. Chrome and Edge (and recent Safari and Firefox) run the model on your GPU through WebGPU, which is the fast path; browsers without WebGPU fall back to running it on the CPU, which works everywhere but takes longer. If a page ever fails to start the engine, updating the browser is usually the fix.

Can I export subtitles?+

Yes. Once the transcript is ready you can copy it, download it as plain text, or download it as SRT or WebVTT subtitles with timestamps — the two formats video editors, YouTube, and media players expect. Click any timestamp on the page to play the recording from that point.

The live version

This is the transcript after the call. Assistly helps during it.

The desktop app follows your meetings, sales calls, and interviews as they happen — the answer is on your screen while the question is still being asked, every speaker kept straight, and the recap with action items is waiting when you hang up. No bot joins the call.