Auto-Generate Captions with On-Device AI Voice Recognition

Learn how to use Recorded's AI Voice Caption feature to automatically transcribe your narration into editable, styled captions—no cloud upload required.

Auto-Generate Captions with On-Device AI Voice Recognition

Typing out captions by hand is one of the most tedious parts of polishing a screen recording. Recorded’s AI Voice Caption feature removes that friction entirely: it listens to your microphone narration and generates timed, editable captions automatically—all processed locally on your machine, with nothing sent to the cloud.

This guide walks through how the feature works, how to get the best results, and how to fit it into your editing workflow.

What AI Voice Caption Does

When you record with your microphone enabled, Recorded can transcribe that narration into caption segments and drop them straight onto your video’s timeline as styled text. Because transcription runs entirely on-device, your audio never leaves your computer—making it a solid option even for narration that includes sensitive or confidential information.

Prerequisites: Record With Microphone Audio

AI Voice Caption transcribes your microphone track, not system audio or on-screen text. If a recording was made without the microphone enabled, the Captions tab will show a notice that no microphone audio is available and generation will be disabled.

Takeaway: if you plan to auto-generate captions, make sure the microphone toggle is on before you start recording.

Step 1: Choose an AI Model

Open your recording in the editor and switch to the Captions tab. You’ll see two model options:

  • Fast — quicker processing, ideal for rough drafts or long recordings
  • Accurate — higher transcription accuracy, better for final, publish-ready captions

The first time you use a given model, Recorded needs to download it before transcription can run. Download progress is shown inline, and once complete the model stays available for future recordings—no repeated downloads.

Step 2: Generate Captions

With a model selected and installed, click Generate Captions. Recorded transcribes your narration in the background and shows progress as it works. When it finishes, you’ll see a list of caption segments, each with a timestamp and the transcribed text.

Step 3: Review and Edit Segments

Auto-generated transcripts are a great starting point, but they’re rarely perfect—homonyms, product names, and technical jargon are common places where speech recognition slips up. Before applying captions to your video:

  • Click any segment to edit its text inline
  • Delete segments you don’t want (for example, filler words or asides you’d rather cut)
  • Watch for segments marked Excluded—these fall inside sections of the video you’ve already trimmed out, so Recorded automatically leaves them out of the final result

Reviewing every segment before applying takes a few minutes but makes a noticeable difference in the finished video’s professionalism.

Step 4: Choose an Apply Mode

Once you’re happy with the captions, decide how they should interact with any text you’ve already added to the timeline:

  • Replace — removes all existing text overlays and applies only the new captions
  • Add — keeps your existing text overlays and adds the captions alongside them

If you’ve already added titles or callouts with the Text tab, Add is usually the safer choice.

Step 5: Style Your Captions

Once applied, captions become normal text segments on your timeline—so you can restyle them exactly like any other text overlay. By default, captions use a clean, high-contrast preset (white text on a semi-transparent dark background, centered near the bottom of the frame), but you’re free to adjust font, size, color, position, and background from the Text tab to match your brand.

Tips for Cleaner Auto-Captions

Since caption quality depends directly on your narration, a bit of prep goes a long way:

  1. Speak clearly and at a steady pace. Rushed or mumbled narration is the most common cause of transcription errors.
  2. Reduce background noise. A quiet room and a decent microphone dramatically improve accuracy.
  3. Pause between thoughts. Natural pauses help the model segment your captions at sensible points.
  4. Use the Accurate model for final exports. Reserve the Fast model for quick previews while you’re still editing.
  5. Trim first, caption second. Generating captions after you’ve already cut out mistakes and dead air means fewer segments to review and discard.

Why On-Device Processing Matters

Because transcription happens locally, AI Voice Caption works without an internet connection and keeps your narration private—useful for internal training videos, customer walkthroughs with sensitive account details, or any recording you’d rather not upload to a third-party service just to get a transcript.

Conclusion

Recorded’s AI Voice Caption feature turns what used to be a manual, time-consuming step into a quick generate-review-apply workflow. Combined with the Text tab’s styling options, it’s an easy way to make your screen recordings more accessible and more engaging—without ever leaving the editor.

Happy recording!