Initializing secure environment…
Initializing secure environment…
Transcribe a recording into a formatted PDF transcript. Speech recognition runs on your own device with Whisper, so a confidential meeting, a medical consultation or a client call is never sent anywhere. Speaker turns and pauses become paragraph breaks, timestamps are optional, and the transcript is laid out as a readable document rather than a wall of text. The recognition model is downloaded once and the size is stated up front. No sign-up, no watermark, no upload.
To convert audio to a PDF transcript free, add an MP3, WAV or M4A file and on-device speech recognition transcribes it into a formatted PDF. Nothing is uploaded. The model is downloaded once on first use, and the size is shown before it starts.
Small is fine for a quiet one-to-one conversation and is quick to download. Medium is the sensible default for a meeting with several speakers. Large is worth the download for anything where the words themselves matter — a deposition, an interview, a dictation you will rely on later.
Ten seconds of microphone placement beats any model. A device 15 to 20 cm from the speaker, in a quiet room, with no fan noise, will be transcribed more accurately than a pristine recording of two people in a noisy café. Post-processing cannot recover what a bad microphone never captured.
On-device speech recognition is not a court reporter and is not certified. It is excellent for notes, minutes and search, and it is not the right tool for a verbatim legal transcript, where a certified transcription service is the appropriate expense. Know which of the two you need before you rely on it.
It works for a short memo and it is slow for anything else, because the model download and the compute are both significant on a phone. Record on the phone, then transcribe on a desktop where the model is already cached. The audio never leaves the device either way.
No. Whisper runs in your browser as WebAssembly, and the audio is decoded and transcribed on your device. The only download is the recognition model — a few hundred megabytes depending on the size you choose, stated up front — and it contains nothing about your recording.
This is the honest trade-off of on-device recognition. A small model is fast and compact with lower accuracy; a large model is far more accurate and substantially bigger. The download size is displayed before it starts, and it is cached afterwards, so the cost is once.
Very good on clear speech in a quiet room with one speaker. Accuracy falls with background noise, overlapping speakers, heavy accents and a poor microphone. Whisper inserts punctuation and casing itself, which is usually right and occasionally creative — read it before sending it to anyone.
It can separate turns where there is a clear pause or a change of voice, and the transcript reflects that as paragraph breaks. It does not perform speaker diarration in the research sense, so two similar voices may merge. Do not rely on it for a transcript where who said what matters legally.
An hour of clean audio is comfortable on a desktop. Longer recordings and phone hardware work but are slow, because the model runs on your CPU or GPU rather than a datacentre. Split a long recording if you want progress rather than a wait.
MP3, WAV and M4A, which covers phone voice memos, meeting recorders and podcast exports. Anything your browser can decode as audio is accepted.
Yes, and you should — proofreading takes minutes and removes the errors that make a transcript embarrassing to send. Corrections are made in the review pane and the PDF is generated from the corrected text.
Yes. Free, no account, no watermark, and no upload of your audio. The model is fetched once and then works offline.
More convert to pdf — all free, no upload.