Tips & tricks

Get the most out of DijiFlow Dictate

A few small tweaks make a big difference to accuracy, speed, and comfort. Here are our favourite ways to fine-tune the app for your voice, your language, and your machine.

Your microphone matters more than any single setting in the app. A clip-on lavalier mic — the little clip that attaches to your shirt or collar — is one of the cheapest and biggest upgrades you can make to accuracy. It sits close to your mouth, keeps a steady distance as you move around, and rejects a lot of the room noise that trips up a laptop's built-in mic.

In short: almost any external mic that gets closer to your mouth than your laptop's built-in one will noticeably improve your results.

The silence threshold is what tells DijiFlow the difference between you speaking and background noise. Getting it right for your room is the single tweak that prevents most dropped words and false triggers. Use the built-in mic test to dial it in.

  1. 1 Open Settings and click Test the mic.
  2. 2 Speak normally and watch the level meter — while you speak, it should sit roughly between the middle and the right of the meter.
  3. 3 Now stay silent — the level should drop and stay below the white line. Where that line needs to sit depends on how noisy your room is.
  4. 4 Drag the silence filter to move the white line up or down. That white line is your silence threshold — everything below it is treated as silence.
  5. 5 Drag the microphone sensitivity to change how high your voice reads on the meter, so your speech comfortably clears the line.

Aim for a clear gap: your speech well above the white line, your silence well below it. Re-check it whenever you change rooms or microphones.

DijiFlow can learn the words that are specific to you — names, jargon, brands — and fix them automatically. Because this happens after transcription, it sharpens the result without ever biasing what the model actually hears.

  1. 1 Open the Vocabulary tab.
  2. 2 Tick the fields that apply to you, so the right term packs are loaded.
  3. 3 Add your own custom terms — colleagues' names, your company, product names, and any words you say often.
  4. 4 If the app keeps mishearing one specific word, add the correct spelling to your custom terms and it will start getting it right.

Think of your custom terms as a running list: every time something comes out wrong, add the correct version and it stops happening.

Bigger models are more accurate, but they are slower and use more memory — RAM on Mac, VRAM on Windows. There's a sweet spot for every machine, so it's worth downloading a couple and testing which one keeps up with you.

  1. 1 For English, small or medium is usually the sweet spot. On lower-spec machines, base still works well.
  2. 2 For other languages or mixed/multilingual use, we strongly recommend medium or large — smaller models lose accuracy fast outside English.
  3. 3 Download a few sizes and test them on your own speech to see what your machine handles comfortably.

If a bigger model feels sluggish, drop one size — a fast model you actually use beats an accurate one you avoid.

Push-to-talk should feel effortless — something you can hold without thinking. The best shortcut is one your fingers already rest near, and you can even move it off the keyboard entirely.

  1. 1 Great keyboard combos to try: Ctrl+Space, Alt+Space, or Shift+Space (especially on Windows).
  2. 2 On Mac, Cmd+Space is taken by Spotlight — either reassign Spotlight in System Settings → Keyboard → Keyboard Shortcuts → Spotlight and use Cmd+Space, or simply use Shift+Space or Ctrl+Space.
  3. 3 Power move: bind push-to-talk and Enter / new-line to your mouse's back and forward buttons — for example Shift+Back for a new line and Shift+Forward for Enter (or bind them directly in your mouse software).

Set the mouse bindings and you can dictate almost entirely with the mouse — no keyboard needed. On Windows use your mouse vendor's utility; on Mac a remapping tool like the ones for gaming mice works well.

Speech recognition is only as good as what reaches the microphone. A fan blowing straight into the mic, a crowded room, or music playing nearby will all pull accuracy down — often more than people expect.

If the noise is unavoidable, raise your silence threshold to match the environment so the app isn't constantly triggered by the background — then lower it again when you're back somewhere quiet.

Dynamic Memory Optimization is the setting to reach for when you're short on RAM (Mac) or VRAM (Windows). It loads and offloads models as you switch, so only one model is ever active at a time — which keeps a lighter machine responsive.

  1. 1 If models feel slow or your machine is tight on memory, turn on Dynamic Memory Optimization in Settings.
  2. 2 From then on, DijiFlow keeps just the model you're using in memory and offloads the others between toggles.

Heads-up: running a file transcription AND live dictation at the same time will slow down if memory is tight — but with enough RAM/VRAM it's perfectly smooth.

If your machine has headroom — roughly more than 8 GB of RAM or 12 GB of VRAM — you can keep two models loaded and switch between them instantly with different hotkeys. It's one of the most productive things power users do.

  1. 1 Bilingual? Bind English to one hotkey and your other language to a second — switch languages without changing a setting.
  2. 2 Single language? Bind two model sizes — a fast one for quick dictation and a more accurate one for when you can wait a moment.
  3. 3 With Q mode you can even mix languages within a single sentence once you get used to it.

This shines when you write an email in one language while chatting or coding in another — all in the same session, no fiddling.

For meetings, accuracy matters more than speed, so pick the most accurate model you can — even if it's slower. Keep in mind that automatic transcription and speaker identification are never 100% perfect and may not capture a whole meeting cleanly on their own. The trick is to hand the transcript to a large AI model to consolidate it into something useful.

  1. 1 Transcribe the meeting with speaker identification, using the best model you have.
  2. 2 If you know how many people spoke, set the speaker count — it makes identification much more accurate (see the speaker-count tip).
  3. 3 Paste the finished transcript into a large AI model — ChatGPT, Gemini, Claude, DeepSeek, or Kimi — and use the ready-made prompt below to get a clean, consolidated summary.
Ready-to-paste meeting prompt
You are helping me turn a raw meeting recording into something I can actually use.

Below is a transcript that was generated automatically from audio or video, including automatic speaker identification. Two things to keep in mind as you read it:
- The speaker labels ("Speaker 1", "Speaker 2", and so on) and the occasional word may be wrong. Where a label or a word clearly conflicts with the context of the conversation, trust the context and quietly correct it.
- You are usually good at working out who is who from the conversation itself, so infer names where you can.

Optional — if I already know who is who, I'll map them here (leave blank otherwise):
- Speaker 1 =
- Speaker 2 =
- Speaker 3 =

Please read the whole transcript and give me a clear, well-organised summary with these sections:
1. Overview — two or three sentences on what the meeting was about and the overall outcome.
2. Key discussion points — the main topics, grouped logically, with the important details under each.
3. Decisions — what was actually agreed (and by whom, if it's clear).
4. Action items — grouped by person: what they own, and any deadline that was mentioned.
5. Blockers & risks — anything raised as a problem, dependency, or concern.
6. Open questions — anything left unresolved or that needs a follow-up.

Keep it concise and easy to scan. Prefer short bullet points over long paragraphs. If something is unclear or missing from the transcript, say so plainly instead of guessing.

Transcript:
[paste your transcript here]

Good to know: you can keep dictating in other apps while a file transcribes in the background.

DijiFlow isn't just for your own voice — it can transcribe any video or audio file you drop into it, complete with speaker identification. That makes it easy to turn a talk, interview, or lecture into text.

  1. 1 Download the video with any online YouTube downloader.
  2. 2 Drag and drop the downloaded file into DijiFlow as a video — it handles the rest.

The same drag-and-drop works for any local audio or video file, not just YouTube.

If your machine has an Intel or AMD GPU, make sure DijiFlow is actually using it — hardware acceleration is dramatically faster than running on the CPU.

  1. 1 Open the GPU tab.
  2. 2 Make sure Intel or AMD acceleration is selected and shows as working.

Warning for Intel GPUs: they're usually built into the CPU and share limited memory. If you want speed on an Intel GPU, load smaller models or enable Dynamic Memory Optimization so the models fit.

When you transcribe a meeting or a file with several people, speaker identification has to guess how many distinct voices there are. You can help it enormously by telling it up front.

If you know exactly how many people speak in the recording, set that number before you transcribe — speaker identification becomes much more accurate when it isn't guessing.