YOUR WORDS, A LITTLE CLEARER

From spoken words
to something you can keep.

Turn English audio or video into an editable transcript and subtitles. Right here, on your device.

Free to useFiles stay on your deviceNo account needed
TRANSCRIPT STUDIOBETA
1Add a file2Transcribe3Edit & export

Drop your audio or video here

Or choose a file from your device

MP3, WAV, M4A, MP4, WebM and moreUp to 50 MB · 5 minutes · English speech
Files stay on your device

Format support depends on your browser and the file's audio encoding.

Your recording stays with you.

Speech recognition runs in your browser. Your audio, video, and transcript are never sent to a transcription server.

The first transcription downloads a speech model. An internet connection is needed to prepare it.
A SIMPLE WORKFLOW

A few steps. All your words.

1

Bring your recording

Choose an English audio or video file from your device.

2

Let it transcribe

Keep this tab open while the model works locally.

3

Make it yours

Review each line, make corrections, and download TXT or SRT.

GOOD TO KNOW

Before you press play

A few things worth knowing about local transcription.

Is CaptionLeaf really free?

Yes. This version has no accounts, payments or transcription credits. It uses your device to process recordings, with an initial limit of 50 MB and 5 minutes.

Does my recording leave my device?

No. Your file is read in this browser and the transcript stays here. The site and speech model are downloaded over the internet; your recording isn't uploaded for transcription.

Which languages can I transcribe?

This version is designed for English speech. You can switch the interface between English and Chinese, but that does not translate the recording or change its recognition language.

What can I download?

TXT is plain text for your notes. SRT contains your edited text with subtitle timing for a video editor or compatible player. Downloading SRT does not add subtitles to your video image.

Why can the first run take longer?

Your browser first downloads and prepares a speech model. It can cache that model for later use. Transcription speed depends on your device and recording; keep the tab open until it finishes.

Will every video or browser work?

Support depends on the file's audio codec and your browser. Start with a short recording in a recent desktop Chrome or Edge. If decoding fails, try a WAV or MP3 file. Always review the result.