Skip to content

Dictation (voice typing)

Dictation turns speech into text at your cursor. It runs a local Whisper model on your CPU. No GPU, and no audio leaves your machine.

  • Sticky notes and text boxes: the D button on the floating text toolbar above the item.
  • Diagram / shape labels: same floating toolbar.
  • Document nodes: the D button on the format toolbar.
  • The AI chat composer: the 🎙 mic button in the composer bar.

Task lists have no dictation. Voice-memo nodes have their own T transcribe button (see Voice memos).

Click D (or the mic) to start, click again to stop. It’s a toggle: no shortcut, no push-to-talk. While recording, the button pulses red and interim text appears every few seconds; the final transcript is inserted when you stop. Clicking away from the field also stops it.

Only one dictation runs at a time across the app.

The first time you dictate, ReelMarkr downloads the speech package (~150 MB) and the model, with a green on the button. After that it’s instant and offline.

Right-click the D button for the Dictation model menu:

ModelSizeNotes
Tiny75 MBFastest, lower quality
Base140 MBRecommended (default)
Small460 MBSlower, better
Medium1.5 GBSlow, near-pro
Large2.9 GBSlowest, best

The menu marks each as ● active, ✓ downloaded, or ⬇ needs download. This choice is shared with voice-memo transcription.

  • Microphone: Settings ▸ Sound ▸ Microphone, shared by dictation, chat, and voice memos.
  • Manage/remove models: Immersion Canvas Settings ▸ Extras ▸ Voice typing / dictation (Whisper) shows what’s installed and frees the space. It re-installs on next use.

Language is auto-detected.