HOME/DOCUMENTATION/Chapter 3: Audio & Video Transcription Guide
OFFICIAL USER GUIDE

Chapter 3: Audio & Video Transcription Guide

Model selection guide featuring Parakeet v3 TDT, VibeVoice, Whisper, Prompt hints, and speaker grouping.

Chapter 3: Audio & Video Transcription Guide

3.1 Supported Media Formats

Scribis natively decodes all major media formats without external conversion tools:

  • Audio: .mp3, .wav, .m4a, .flac, .aac, .ogg, .wma, .opus
  • Video: .mp4, .mov, .mkv, .avi, .webm, .wmv, .flv

SCREENSHOT BLUEPRINT

Media file queue view displaying imported audio and video files with duration and waveform thumbnails


3.2 ASR Model Selection Guide

Scribis integrates top-tier open-source speech recognition models tailored for diverse workloads:

SCREENSHOT BLUEPRINT

Radar chart comparing ASR models across speed, English accuracy, multilingual accuracy, and CPU efficiency

1. 🦅 Parakeet v3 TDT (#1 Ranked for English Speech ⭐⭐⭐⭐⭐)

  • Highlights: The ultimate model for English transcription. Delivers the lowest Word Error Rate (WER) in modern benchmarks, blazing inference speeds (3x to 16x real-time speed on CPU/GPU), and word-level timestamps.
  • Best For: English podcasts, video essays, webinars, interviews, YouTube subtitles, and academic lectures.

2. 🏆 VibeVoice (#1 Ranked Overall & Multilingual ⭐⭐⭐⭐⭐)

  • Highlights: State-of-the-art conversational voice understanding with zero hallucinations. Excels at complex sentence boundaries, background noise suppression, and nuanced long-form context.
  • Best For: High-stakes interviews, business conferences, documentary transcripts, and multilingual content.

3. ⚡ SenseVoice Small (Ultra-Fast CPU with Emotion & Event Detection)

  • Highlights: Operates at lightning speed on ordinary CPUs with virtually zero memory overhead. Automatically identifies emotions (happy, angry, sad) and audio events (laughter, applause, music).
  • Best For: Short-form videos, quick memos, and live dictation on thin-and-light laptops.

4. 🇹🇼 Breeze-ASR 26 (Specialized for Traditional Chinese & Code-Switching)

  • Highlights: Tuned specifically for Traditional Chinese, Taiwanese accents, Cantonese, Hakka, and mixed Chinese-English speech.

5. 🌐 Whisper Series (Whisper Large-v3 / Distil-Whisper)

  • Highlights: Robust general-purpose transcription with resilient performance in noisy environments.

3.3 Prompt Engineering for Transcription

Use the "Prompt" input in transcription settings to guide the AI model:

  • Inject Technical Terms & Jargon:

    Example Prompt: "A tech talk discussing React 19, TypeScript, Next.js, Kubernetes, Docker, and WebAssembly."

  • Enforce Proper Punctuation & Style:

    Example Prompt: "Formal business English transcription with complete punctuation and capitalized proper nouns."

  • Filter Verbal Fillers:

    Example Prompt: "Omit filler words such as 'um', 'uh', 'you know', and 'like'."


3.4 Speaker Diarization & Auto-Recognizing Speaker Groups

When transcribing interviews, roundtables, or team meetings:

SCREENSHOT BLUEPRINT

Speaker diarization settings and preview with different speaker badges highlighted in distinct colors

  1. Enable "Speaker Diarization".
  2. Set the expected Number of Speakers (e.g., 2 for an interview, or leave -1 for auto-detection).
  3. 👥 Speaker Group Matching (Voiceprint Groups):
    • Once speakers are identified, you can save them into a Group (e.g., "Weekly Product Sync" or "Co-host Team").
    • Next time you transcribe audio from this group, simply select the Group, and Scribis will automatically identify and label each person by name!