Qwen3-TTS vs Kokoro: Which Text-to-Speech Model Is Better for Global Voiceovers?
Compare Qwen3-TTS and Kokoro for multilingual voiceovers, voice design, voice cloning, speed, model size, and licensing. Learn how to create scripts and narration in Scribis.
Scope and product note: Qwen3-TTS and Kokoro are integrated with Scribis. Chatterbox and VoxCPM2 are external comparison references. This article focuses on English-speaking and international content teams creating voiceovers in English, Japanese, Chinese, and other supported languages.
The short answer: Kokoro is lightweight and efficient; Qwen3-TTS offers broader voice control
Qwen3-TTS and Kokoro solve different text-to-speech problems. Kokoro-82M is a small open-weight TTS model designed for efficient, high-quality fixed-voice synthesis. Qwen3-TTS provides Base, CustomVoice, and VoiceDesign paths for multilingual generation, voice cloning, and natural-language control.
Use Kokoro when you need fast, repeatable narration at a low resource cost. Use Qwen3-TTS when you need a multilingual voiceover, a described voice style, or a reference voice. Both are integrated with Scribis, which allows the script, transcript, captions, timeline, and generated audio to live in the same content workflow.

Figure: CV3-Eval vendor-reported snapshot from the IndexTTS project. Lower WER/CER is better; higher speaker similarity is better. It is not a universal TTS ranking.
What is Qwen3-TTS?
Qwen3-TTS is an Apache-2.0 open TTS family supporting Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian. The family contains several useful operating modes.
The Base variants can perform zero-shot voice cloning from a short reference clip. CustomVoice provides built-in voices with instruction and style control. VoiceDesign allows a user to describe a voice in natural language, including attributes such as age, gender, pitch, speaking rate, and emotional style.
This makes Qwen3-TTS useful for character narration, multilingual training content, localized product videos, and voiceover revisions where the same script must be rendered in different styles. Voice cloning also introduces consent, identity, and usage-rights considerations; an open license for the model does not grant rights to clone a person’s voice.
What is Kokoro-82M?
Kokoro-82M is an 82M-parameter open-weight TTS model released under Apache-2.0. Its strength is the combination of a small model footprint, fast synthesis, and a collection of fixed voices.
Kokoro is a practical choice for batch narration, local applications, previews, and teams that do not need custom voice design. Its small size can make it easier to deploy, but language and accent availability depend on the selected voice and version. A multilingual model label should not be interpreted as every voice supporting every language.
| Dimension | Qwen3-TTS | Kokoro-82M |
|---|---|---|
| Main positioning | Multilingual voice control, cloning, and voice design | Lightweight fixed-voice narration |
| Approximate size | 0.6B and 1.7B variants | 82M |
| Voice control | Base cloning, CustomVoice, VoiceDesign, instructions | Primarily fixed voices |
| Language coverage | 10 major languages in the published positioning | Depends on voice and version |
| Best use case | Multilingual production, characters, style variation | Fast previews and high-volume narration |
| License | Apache-2.0 | Apache-2.0 |
| Scribis status | Integrated | Integrated |
Comparison with Chatterbox and VoxCPM2
Chatterbox is an external open TTS family from Resemble AI. Its multilingual model is reported at roughly 500M parameters with 23 or more languages, while other variants target speed or smaller deployment. The project also documents voice cloning, paralinguistic tags, and an audio watermarking route.
VoxCPM2 is another external comparison model. OpenBMB describes it as a 2B tokenizer-free TTS model supporting 30 languages, voice design, controllable cloning, and 48kHz output. Its README reports RTF values for an RTX 4090 under particular PyTorch and Nano-vLLM setups.
These models illustrate why TTS selection needs multiple dimensions. Qwen3-TTS emphasizes language and voice control; Kokoro emphasizes small-model efficiency; Chatterbox emphasizes expressive multilingual cloning and watermarking; VoxCPM2 emphasizes multilingual output, high sample rate, and its tokenizer-free approach.
How to interpret TTS benchmarks
Text-to-speech evaluation is not the same as ASR evaluation. A generated clip can be converted back to text with an ASR model to estimate intelligibility, while speaker similarity estimates how closely it matches a reference voice. Neither metric fully captures naturalness, prosody, pronunciation, emotion, or listener preference.
An IndexTTS CV3-Eval table reports IndexTTS2.5 at 6.75% average WER and 73.18% speaker similarity, VoxCPM2 at 7.22% and 72.02%, and OmniVoice at 6.76% and 70.39%. The table is a vendor-reported project snapshot with model- and language-specific conditions; it should not be treated as a universal Qwen3-TTS or Kokoro ranking.
For an international voiceover team, build a human evaluation set with English names, acronyms, numbers, Japanese names, Chinese proper nouns, long sentences, different speech rates, and emotional instructions. Ask native or highly fluent listeners to rate pronunciation, naturalness, pacing, and speaker consistency.
A practical Scribis workflow
Start with a transcript or script in Scribis. Use Kokoro to generate a fast narration preview and Qwen3-TTS when you need a particular style, a multilingual version, or a reference voice. Review the generated audio alongside the caption timeline and script so that pronunciation fixes can be applied consistently.
This workflow is especially useful when a video changes late in production. If the script is edited, the team can update the text, regenerate the affected narration, and keep the voiceover aligned with the captions. The operational advantage is not only audio quality; it is reduced friction between text, timing, and publishing.
FAQ
Is Qwen3-TTS more natural than Kokoro?
There is no universal answer. Qwen3-TTS provides more voice-control modes, while Kokoro emphasizes a small model and fixed voices. Naturalness depends on language, voice, prompt, text, and listener preference.
Can Kokoro generate every global language?
No. The actual language and accent options depend on the selected voice and version. Check the voice card and the Scribis interface rather than assuming that a multilingual package supports every language equally.
Can Qwen3-TTS clone any person’s voice?
It provides a reference-audio voice-cloning route, but you must obtain consent and verify the legal and commercial rights for any voice used. Model licensing does not replace speaker permission.
Which model is best for a YouTube voiceover?
Kokoro is a good first test for fast fixed-voice narration. Qwen3-TTS is a better candidate when a video needs multiple languages, voice design, or a consistent reference voice. Evaluate the rendered audio in the final video timeline.
Conclusion: choose speed for previews and control for final voiceovers
Kokoro-82M is a practical small model for fast, repeatable narration. Qwen3-TTS is a more flexible choice for multilingual voiceovers, style instructions, voice design, and cloning. Chatterbox and VoxCPM2 provide useful external comparisons. For content teams, Scribis connects the TTS decision to the script, captions, timeline, and final export.
CTA: Create and compare multilingual voiceovers with Qwen3-TTS and Kokoro in Scribis, then review the generated audio in the context of your final video or caption timeline.