Chapter 3: Audio & Video Transcription Guide
Model selection guide featuring Parakeet v3 TDT, VibeVoice, Whisper, Prompt hints, and speaker grouping.
Chapter 3: Audio & Video Transcription Guide
3.1 Supported Media Formats
Scribis natively decodes all major media formats without external conversion tools:
- Audio:
.mp3,.wav,.m4a,.flac,.aac,.ogg,.wma,.opus - Video:
.mp4,.mov,.mkv,.avi,.webm,.wmv,.flv
Media file queue view displaying imported audio and video files with duration and waveform thumbnails
3.2 ASR Model Selection Guide
Scribis integrates top-tier open-source speech recognition models tailored for diverse workloads:
Radar chart comparing ASR models across speed, English accuracy, multilingual accuracy, and CPU efficiency
1. 🦅 Parakeet v3 TDT (#1 Ranked for English Speech ⭐⭐⭐⭐⭐)
- Highlights: The ultimate model for English transcription. Delivers the lowest Word Error Rate (WER) in modern benchmarks, blazing inference speeds (3x to 16x real-time speed on CPU/GPU), and word-level timestamps.
- Best For: English podcasts, video essays, webinars, interviews, YouTube subtitles, and academic lectures.
2. 🏆 VibeVoice (#1 Ranked Overall & Multilingual ⭐⭐⭐⭐⭐)
- Highlights: State-of-the-art conversational voice understanding with zero hallucinations. Excels at complex sentence boundaries, background noise suppression, and nuanced long-form context.
- Best For: High-stakes interviews, business conferences, documentary transcripts, and multilingual content.
3. ⚡ SenseVoice Small (Ultra-Fast CPU with Emotion & Event Detection)
- Highlights: Operates at lightning speed on ordinary CPUs with virtually zero memory overhead. Automatically identifies emotions (happy, angry, sad) and audio events (laughter, applause, music).
- Best For: Short-form videos, quick memos, and live dictation on thin-and-light laptops.
4. 🇹🇼 Breeze-ASR 26 (Specialized for Traditional Chinese & Code-Switching)
- Highlights: Tuned specifically for Traditional Chinese, Taiwanese accents, Cantonese, Hakka, and mixed Chinese-English speech.
5. 🌐 Whisper Series (Whisper Large-v3 / Distil-Whisper)
- Highlights: Robust general-purpose transcription with resilient performance in noisy environments.
3.3 Prompt Engineering for Transcription
Use the "Prompt" input in transcription settings to guide the AI model:
- Inject Technical Terms & Jargon:
Example Prompt: "A tech talk discussing React 19, TypeScript, Next.js, Kubernetes, Docker, and WebAssembly."
- Enforce Proper Punctuation & Style:
Example Prompt: "Formal business English transcription with complete punctuation and capitalized proper nouns."
- Filter Verbal Fillers:
Example Prompt: "Omit filler words such as 'um', 'uh', 'you know', and 'like'."
3.4 Speaker Diarization & Auto-Recognizing Speaker Groups
When transcribing interviews, roundtables, or team meetings:
Speaker diarization settings and preview with different speaker badges highlighted in distinct colors
- Enable "Speaker Diarization".
- Set the expected Number of Speakers (e.g.,
2for an interview, or leave-1for auto-detection). - 👥 Speaker Group Matching (Voiceprint Groups):
- Once speakers are identified, you can save them into a Group (e.g., "Weekly Product Sync" or "Co-host Team").
- Next time you transcribe audio from this group, simply select the Group, and Scribis will automatically identify and label each person by name!