Tutorial

Real-time Speaker Diarization with Nemotron-3 in Scribis

personJH LAI
calendar_today

Learn how to use NVIDIA Nemotron-3 Diarization in Scribis to identify speakers in real time during live recording and jump to speaker segments on the timeline.

In meetings, interviews, podcasts, or group discussions, live speech-to-text alone often falls short. If transcript text from multiple speakers is merged into a single block, you still have to spend extra time figuring out who said what.

Scribis integrates with NVIDIA Nemotron-3 Diarization to identify distinct speakers during live recording and transcription, tagging each turn with color-coded Speaker labels in real time.

Scribis AI transcription and real-time speaker diarization interface

What You'll Learn

In this tutorial, you will learn how to:

  • Download the Nemotron-3 Diarization model
  • Enable live transcription and Progressive Updates
  • Display real-time speaker labels while recording
  • Use the timeline in Project Workspace to jump instantly to specific speaker audio segments

Scribis supports both macOS and Windows. While screenshots and examples use macOS, speaker diarization functionality works seamlessly across desktop platforms.


What is Speaker Diarization?

Speaker Diarization—often called speaker separation or speaker identification—answers a fundamental question:

Who spoke when?

Standard speech-to-text transcribes audio into text sequentially:

Welcome back to the deep dive. Oh yeah. I have to ask...

With Speaker Diarization enabled, Scribis structures the transcript clearly:

Speaker 1

Welcome back to the deep dive...

Speaker 2

Oh yeah.

Speaker 1

I have to ask...

This distinction is essential for multi-speaker meetings, interviews, and podcast production.


Step-by-Step Guide

Step 1: Download Nemotron-3 Diarization

Open Scribis Settings and navigate to:

Models → Separate Speakers

Locate NVIDIA Nemotron-3 Diarization and click download.

The Nemotron-3 Diarization model weighs only 107 MB. You don't need to download gigabytes of heavy models to unlock local speaker diarization capability on your machine.

Step 2: Open the Floating Recording Bar

Launch the Floating Recording Bar from the top bar in Scribis.

The floating bar lets you control recording and transcription while using other applications, watching videos, or hosting video calls. From the floating bar, you can:

  • Start recording
  • Pause recording
  • Stop recording
  • Monitor recording duration
  • Check real-time audio input levels

Step 3: Enable Transcription and Progressive Updates

Before starting a recording, confirm that the following preferences are enabled:

  • Transcription
  • Progressive Updates

Transcription converts incoming audio into text on the fly, while Progressive Updates continuously refreshes the transcript stream instead of waiting for audio processing to finish.

With both settings active, your transcript streams seamlessly in near real-time.

Step 4: Start Recording

Click the Record button. Play external media or start your live session:

  • Meetings & Conferences
  • Interviews
  • Podcasts
  • Lectures & Classes
  • Press Conferences
  • Group Discussions

Scribis continuously captures incoming audio and generates live transcripts.

Step 5: View Real-time Speaker Identification

As Scribis detects speaker changes, it automatically attaches color-coded Speaker tags to transcript chunks:

Speaker 1

Welcome back to the deep dive. You know, I had a bit of a moment this morning...

Speaker 2

Oh yeah.

Speaker 1

A full verse chorus, the whole thing just sort of landed in my head...

Speaker 2

I hate to ruin the magic for you right at the start...

Distinct speaker tags make it easy to tell who is talking at a glance while recording is in progress.

This highlights the main advantage over standard Live Captions: standard live captions convert Audio → Text, whereas Diarization delivers Audio → Text → Speaker. You learn not only what was said, but who said it.

Step 6: Open Project Workspace After Recording

Once recording stops, open the Project Workspace.

Real-time transcriptions, timestamps, and speaker metadata are preserved inside your project. You can review the complete transcript alongside timeline waveforms.

Step 7: Jump to Speaker Segments via Timeline

Speaker Diarization extends beyond transcript tags. Scribis plots speaker turns directly onto the visual audio timeline.

Clicking any Speaker Interval instantly seeks playback to that exact audio timestamp. If you need to re-verify a comment from Speaker 2, you don't need to scrub through the entire recording manually:

Click Speaker Segment → Jump to Timestamp → Play

For long meeting recordings, podcasts, or interviews, timeline navigation saves substantial review time.


Complete End-to-End Workflow

Models → Separate Speakers (Nemotron-3 Diarization)
                        ↓
             Open Floating Recording Bar
                        ↓
    Enable Transcription & Progressive Updates
                        ↓
                    Start Recording
                        ↓
            Real-time Speaker Labels
                        ↓
            Open Project Workspace
                        ↓
    Click Speaker Interval on Timeline to Jump & Play

Key Concepts & FAQ

Why Choose Nemotron-3 Diarization?

Scribis includes Nemotron-3 Diarization as a compact, high-efficiency speaker diarization option. At just 107 MB, it integrates cleanly into local desktop workflows without heavy memory overhead.

It is particularly effective for:

  • Live subtitles & captions
  • Meeting minutes & notes
  • Podcast editing
  • Interview transcripts
  • Multi-person audio recordings
  • YouTube and video transcription

Nemotron-3 runs entirely on-device, meaning speaker diarization does not require sending raw audio to external servers.

Live Captions vs. Speaker Diarization

Without diarization:

Hello everyone. Hi, thanks for having me. Let's start with today's topic.

With speaker diarization: Speaker 1

Hello everyone.

Speaker 2

Hi, thanks for having me.

Speaker 1

Let's start with today's topic.

For single-speaker audio, live captions work fine. But for conversations involving two or more participants, diarization provides critical structure.

Recommended Use Cases

  • Meeting Records: Instantly attribute decisions and action items to specific attendees.
  • Podcasts: Establish clear speaker structures right during the initial recording stage.
  • Interviews: Automatically separate interviewer and interviewee dialogue without manual editing.
  • Lectures & Discussions: Distinguish teacher instructions from student questions.
  • Video Content: Create structured transcripts and subtitles while watching external videos or webinars.

Summary

Setting up real-time speaker diarization in Scribis takes just a few steps:

  1. Navigate to Models → Separate Speakers
  2. Download Nemotron-3 Diarization
  3. Open the Floating Recording Bar
  4. Enable Transcription and Progressive Updates
  5. Start recording
  6. Monitor real-time Speaker 1, Speaker 2 labels
  7. Jump directly to speaker segments on the Project Workspace timeline after recording

With Nemotron-3 Diarization, Scribis transforms live transcription from simply "what was said?" into "who said what, and when?".


References

  1. Scribis Official Website
  2. NVIDIA Nemotron Catalog