Local Transcriber

100% FreeNo sign-upRuns on your device

Private speech-to-text, processed on your device. Press Start, speak, and watch the transcript appear — then download it as Word, text or subtitles.

Allow the microphone to transcribe live speech

Your browser will ask for permission. Audio is turned into text on this device and is not sent anywhere.

NOT STARTED

0s

Press Prepare to load the speech model.

Microphone
Off
Processing
On this device
Waiting
Nothing
Last text
—
Saved locally
Not yet

Speech model

Both models are stored on this device and neither is downloaded from anywhere else. Changing model reloads the engine, so it can only be done between sessions.

Transcript

0 words

Privacy and saved transcripts

Your speech stays on your device

  • Microphone audio is turned into text here, in your browser.
  • Your audio and your transcript are not sent to a transcription server.
  • There is no account, and nothing is synced anywhere.
  • Transcripts are stored in this browser, on this device, until you delete them.
  • The page and the speech model are downloaded from Zerdly, like any web page. After that, transcription needs no network.
  • Audio is released as soon as it has been transcribed — no recording is kept.
  • Clearing site data for Zerdly in your browser removes saved transcripts and the offline model. Files you have exported are unaffected.

Saved on this device

Checking…

Removing the speech model does not delete transcripts, and deleting transcripts does not remove the model.

How it works

A speech-recognition model is downloaded to your browser once, and after that every word is recognised on your own computer. Your microphone audio does not travel to a transcription service, and neither does the text it produces.

Audio is handled in short windows: a few seconds are captured, turned into text, and released. That is what allows a session to run for hours without the tool needing more and more memory as the transcript grows.

Long meetings and interviews

This is built for long-form speech rather than short voice notes. Each finished piece of transcript is written to your device as soon as it exists, so an accidental refresh or a browser crash does not cost you the session — you are offered it back when you return.

A desktop or laptop browser is the right place for a multi-hour session. Phones and tablets work, but they suspend background tabs and throttle heavy work, and local speech recognition is heavy.

Working offline

Open Offline use and choose Make available offline. That stores the speech model and the runtime on your device, and the tool will then transcribe with no internet connection at all.

The badge only says Offline ready once every file has been verified as present — not merely because something was downloaded once.

Correcting the transcript

Nothing recognises speech perfectly. Click any line to correct it, and the original wording is always kept underneath, so Restore original is available afterwards.

For a word the model mis-hears every single time — a place name, a piece of software, a surname — Replace fixes every occurrence at once and tells you how many it changed.

Exporting

Word documents, plain text, Markdown, SRT and WebVTT subtitles, and a JSON backup of the session. Timestamps, session details and markers are optional, so the Word file you hand to somebody reads like a document rather than a log.

Everything is built in your browser and saved straight to your downloads. No file is generated on a server.

What it does not do

It does not keep a recording of your audio — sound exists only for as long as it takes to transcribe, then it is released. It does not identify who is speaking. It does not claim an accuracy figure, because an honest one depends on your microphone and your room.

And it is not a substitute for asking permission: make sure you have any consent you need before transcribing or recording other people.

Frequently asked questions

Is my speech uploaded anywhere?
No. The speech model runs inside your browser, so your microphone audio is turned into text on your own device and is not sent to a transcription server. The model file itself is downloaded from Zerdly once, the same way any web page is downloaded — after that, transcription needs no network at all.
Do I need an account?
No. There is no sign-up, no login and no cloud sync. Your transcripts are stored in your browser on this device, and you can delete them at any time.
How accurate is it?
It varies with your microphone, the room and how clearly people speak, so no honest figure can be quoted. Expect to proofread names, technical terms and numbers. You can correct any line, and Replace all will fix a word the model gets wrong every time.
Can I choose a more accurate model?
Yes. There are two, and you pick before you start: a faster one (about 41 MB) and a more accurate one (about 75 MB) that handles names, technical terms and accents better. Both are stored on this device and neither is downloaded from anywhere but Zerdly. The faster one is the default because it keeps up comfortably on most machines — the larger one asks noticeably more of your device, and on a slower computer the transcript can fall behind while you are still speaking. If that happens the tool says so and offers to finish and switch.
How long can a session be?
The design target is long sessions — meetings and interviews rather than voice notes. Audio is processed in short windows and released immediately, so a transcript growing for hours does not make the engine use more memory. Five-hour continuous runs have been measured, transcribing over 46,000 words without the tool slowing down or losing its place. Desktop or laptop is the intended use.
Does it work offline?
Yes, once you have installed the speech model. Open the Offline use section and choose Make available offline; after that the tool can transcribe with no internet connection. You need to be online once, for the initial download.
What happens if my browser crashes or I refresh by accident?
Each finished piece of transcript is saved to your device as it is produced, not at the end. If the page goes away, you are offered your session back when you return. Anything spoken in the last few seconds before the interruption was not transcribed yet and cannot be recovered — the tool says so rather than implying otherwise.
Which formats can I export?
Word (.docx), plain text, Markdown, SRT and WebVTT subtitles, and a JSON backup of the whole session. Every file is built in your browser and saved straight to your downloads.
Is audio recorded and kept?
Only if you ask for it. By default audio exists just for the moment it takes to transcribe it and is then released, so nothing you say is stored. Before you start you can tick "Also keep the audio", which saves a recording on this device — roughly 14 MB an hour — that you can play back, download, or delete on its own without touching the transcript. It is never uploaded, and the choice is made fresh each session rather than remembered.
Where is a recording kept, and how do I get rid of it?
In your browser's storage on this device, alongside the transcript and nothing else. Deleting the session deletes its recording with it, and the privacy panel can remove a recording on its own if you want to keep the words but not your voice.
What do I need for it to work?
A current desktop browser is the best experience, because local speech recognition is demanding. The tool checks your browser when it loads and tells you plainly if something it needs is missing, rather than failing silently.

More tools that run on your device