HOME/DOCUMENTATION/Chapter 1: Introduction & Quick Start
OFFICIAL USER GUIDE

Chapter 1: Introduction & Quick Start

Discover Scribis local-first privacy, cutting-edge ASR models like Parakeet & VibeVoice, and transcribe your first file in 3 minutes.

Chapter 1: Introduction & Quick Start

1.1 What is Scribis?

Scribis is a next-generation AI speech-to-text, subtitle editor, and multimodal creation studio crafted for video creators, translators, researchers, podcasters, and professionals.

SCREENSHOT BLUEPRINT

Scribis desktop main user interface showing navigation sidebar, quick action cards, and recent projects

Unlike traditional cloud-based transcription tools that charge per minute and upload sensitive audio to remote servers, Scribis is built with a Local-First Privacy architecture. Your audio files and transcripts remain 100% on your device unless you explicitly connect your own cloud AI keys.

🌟 Key Product Highlights

  1. 🔒 100% On-Device Privacy: Speech recognition and media rendering run directly on your GPU/CPU with zero cloud leakage.
  2. Top-Tier Open-Source ASR Models: Features world-class models including Parakeet v3 TDT (#1 for English with ultra-fast speed), VibeVoice (state-of-the-art context and zero hallucinations), SenseVoice Small (ultra-fast CPU inference with emotion detection), and Whisper Large-v3.
  3. 🎬 Text-Driven Video Editing: Edit video by editing text. Jump playheads by clicking any word, perform rough cuts by deleting subtitle lines, and add dynamic zoom-in pans, privacy mosaics, and title overlays directly on the timeline.
  4. 🎙️ Global Dictation & Mini Recorder: Press a global hotkey anywhere on your desktop to dictate smoothly with smart auto-formatting, grammar polishing, and instant paste into Word, Notion, IDEs, and messengers.
  5. 🤖 AI Assistant & GEC Correction Sync: Chat with lengthy recordings, generate structured meeting summaries, and perform homophone error correction with one-click sync to your personal Correction Dictionary.

1.2 3-Step Quick Start Guide

Follow these three simple steps to transcribe your first audio or video file in under 3 minutes:

Step 1: Import Media Files

Open Scribis and click "Transcribe" in the left sidebar. Drag and drop your audio or video file (.mp3, .wav, .mp4, .mov, etc.) into the upload zone.

SCREENSHOT BLUEPRINT

Drag and drop media file into the transcription area with waveform and duration preview

[!TIP] You can also click "Record from Mic" to record audio directly in the app.


Step 2: Choose ASR Model & Language

Configure your transcription settings in the right panel:

  1. Model Selection:
    • For English Audio (#1 Recommended): Select Parakeet v3 TDT (fastest inference, lowest Word Error Rate).
    • For Maximum Precision & Multilingual Context (#1 Overall): Select VibeVoice.
    • For Low-Resource Laptops / CPU Inference: Select SenseVoice Small.
  2. Language: Set to Auto or explicitly select your spoken language (e.g., English, Chinese (Traditional), Japanese).
  3. Speaker Diarization: Turn on Speaker Diarization if your media contains multiple speakers.

SCREENSHOT BLUEPRINT

Transcription configuration panel with model dropdown, language selector, prompt field, and speaker diarization switch

Click "Start Transcription".


Step 3: Review, Edit & Export

Once transcription finishes, Scribis seamlessly opens the Translate & Subtitle Editor:

  1. Interactive Review: Click any subtitle line to jump video playback to that exact timestamp. Edit words directly or use shortcuts (Alt + E edit, Alt + M merge, Alt + D delete).
  2. AI Error Correction (GEC): Click "GEC Correction" to inspect suggested typo fixes in a red/green Diff Review panel and sync accepted fixes to your dictionary.
  3. Export: Click "Export" in the top right corner to download .srt, .vtt, .txt, .fcpxml, or render a hardcoded burn-in MP4 video.

SCREENSHOT BLUEPRINT

Subtitle editor workspace showing video preview at top, subtitle list in middle, and audio waveform timeline at bottom