Free tools

A free audio-to-text tool that runs in your browser

By The Assistly team ·

Most "free" transcription tools are a sign-up form with a stopwatch behind it. Ten free minutes, then a plan; or a transcript you can read on screen but only export once you've paid. We wanted something we could point people to without that asterisk, so we built one: tryassistly.com/tools/audio-to-text.

Drop in an MP3, a WAV, an M4A, the audio track of a video, or hit Record and talk. A minute or so later you have a transcript with timestamps, which you can copy, download as plain text, or download as SRT or WebVTT subtitles. No account, no watermark, no export gate.

Why it can be free

Transcription is priced by the minute because someone is paying for the server your audio runs on. This tool doesn't have one. It downloads an open-source speech model — OpenAI's Whisper, the same family most paid products are built on — into your browser, once, and runs it on your own computer. On a machine with a modern GPU the browser uses it through WebGPU, and a recording transcribes several times faster than real time. Older hardware falls back to the CPU, which is slower but works everywhere.

Because the work happens on your device, the audio never leaves it. There is nothing for us to store, nothing for us to see, and nothing to delete later. The only network request the tool makes is the one that fetches the model, and your browser caches that so the next transcript starts immediately. You can check this yourself: load the model, switch off your Wi-Fi, and keep going.

What you get

  • Three model sizes. Fast for clean speech in a hurry, Balanced for most recordings, Accurate when names, accents, or background noise matter more than the wait.
  • 99 languages, detected automatically or set by hand.
  • Timestamps you can click. Every line replays the recording from that point, which is the fastest way to check a word you don't trust.
  • Exports that are actually free. Plain text, SRT, and WebVTT — the formats video editors, YouTube, and media players expect.

What it doesn't do

It transcribes what was said, not who said it. Whisper has no idea how many people are in the room, so a meeting comes out as one voice with timestamps.

That gap is not an accident; it's where the tool and the product part ways. The free tool is the transcript after the call. Assistly is what happens during one: it follows the conversation as it happens, keeps every speaker straight, puts the answer on your screen while the question is still being asked, and writes the recap with action items when you hang up. If the transcript you just made is from a meeting you wish you'd had help in, that's the desktop app.

For everything else — a voice memo, a lecture, an interview you recorded, a video that needs captions — the free tool is right there, and it stays free.

Ready to put Assistly in your corner?

Real-time guidance in your meetings, calls, and interviews — and clean notes after every session. Free to start, no card required.

Get started for free