Scribis Custom TTS vs Qwen3-TTS and Kokoro: Choosing English, Japanese, and Chinese Voices
Compare Scribis custom TTS with Qwen3-TTS, Kokoro, and Piper for English, Japanese, Chinese, voiceover workflows, accent control, speed, licensing, and production use.
Scope and product note: Scribis custom TTS, Qwen3-TTS, and Kokoro are discussed at different levels. Scribis custom TTS is integrated with Scribis and supports Chinese, Japanese, English, and a Taiwan accent. Qwen3-TTS and Kokoro are also integrated with Scribis. Piper is included as an external local-deployment reference; external model descriptions do not mean that every Piper voice or backend is available in Scribis.
The short answer: start with Scribis when the workflow and supported voices matter most
The best text-to-speech engine is not always the largest model or the one with the highest language count. For an international content team, the important questions are whether English, Japanese, and Chinese sound natural; whether the preferred accent is available; whether scripts can be regenerated quickly; and whether audio stays aligned with captions and video timelines.
Scribis custom TTS supports English, Japanese, Chinese, and a Taiwan accent inside the Scribis content workflow. Qwen3-TTS is an attractive integrated option when you need voice design, voice cloning, or instruction-based style control. Kokoro is a useful lightweight fixed-voice option. Piper is a valuable external reference for local and CPU-oriented deployments, but its engine and voice packs must be checked separately.

Figure: Qualitative feature positioning, not a voice-quality ranking. Language and accent availability for external models depends on the selected voice and version.
Why accent is a separate TTS requirement
English, Japanese, and Chinese support is not the same as natural regional speech. TTS quality also depends on pronunciation, prosody, numbers, acronyms, names, punctuation, speaking rate, and how a voice handles long paragraphs.
For an international team, “English” might mean a neutral narration voice, US English, UK English, Indian English, Singapore English, or another regional style. A model can support English while offering only a limited selection of accents. The same principle applies to Japanese and Chinese. Always evaluate the actual voice and script rather than the language label alone.
A Taiwan accent is one optional regional requirement in the Scribis product matrix. It is important for Taiwan-facing content, but the English edition of this article treats it as an additional selection dimension rather than the default international use case.
Scribis custom TTS: the workflow-first option
Scribis custom TTS is designed as part of the same content workflow that handles transcription, captions, timelines, and editing. A user can start from a transcript or script, revise the text, generate voiceover, and review the audio against the caption timeline without manually moving every file between unrelated tools.
The product-provided capability list confirms support for Chinese, Japanese, English, and a Taiwan accent. Internal architecture, parameter count, and independent MOS or speaker-similarity results are not assumed in this article. It would be misleading to invent a numeric benchmark for a custom product system that has not published one.
For content teams, this means the appropriate evaluation is end-to-end: pronunciation, naturalness, voice consistency, regeneration speed, caption synchronization, and time to final export. A custom TTS product can be operationally better even when an external model has a more detailed public research paper.
Qwen3-TTS: more control for multilingual voice production
Qwen3-TTS is an Apache-2.0 open TTS family that lists Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian among its major supported languages.
Its Base variants support short-reference voice cloning; CustomVoice provides built-in speakers and instruction or style control; VoiceDesign allows a user to describe a voice in natural language. This makes Qwen3-TTS useful for multilingual characters, training content, localized marketing, and projects that need more variation than a fixed narrator.
Voice cloning requires careful governance. Consent from the speaker, allowed usage, disclosure, and commercial rights should be verified before a reference recording is used. An Apache-2.0 model license does not grant permission to imitate a person’s voice.
Kokoro and Piper: lightweight local references
Kokoro-82M is an 82M-parameter open TTS model released under Apache-2.0. It is a strong candidate for fast previews, repeatable narration, and small deployments. The actual language and accent coverage still depends on the selected voice and release.
Piper is commonly used as a local neural TTS engine with community voice packs and CPU-friendly deployment. CrispASR lists Piper community voices as a TTS backend, but the current project and each voice model should be checked for licensing before commercial use. This article therefore treats Piper as an external deployment reference, not as a blanket claim that all Piper voices are available in Scribis or commercially unrestricted.
| Dimension | Scribis custom TTS | Qwen3-TTS | Kokoro-82M | Piper |
|---|---|---|---|---|
| Product status | Integrated in Scribis | Integrated in Scribis | Integrated in Scribis | External local ecosystem reference |
| English | Supported | Supported | Depends on voice | Depends on voice pack |
| Japanese | Supported | Supported | Depends on voice | Depends on voice pack |
| Chinese | Supported | Supported | Depends on voice | Depends on voice pack |
| Taiwan accent | Supported | Test the selected voice | Test the selected voice | Depends on voice pack |
| Voice control | Defined by Scribis product capabilities | Cloning, CustomVoice, VoiceDesign | Primarily fixed voices | Primarily fixed voices |
| Main strength | Integrated script-to-voice content workflow | Multilingual voice control | Small model and fast synthesis | Local and edge deployment |
How to benchmark multilingual TTS
TTS needs at least three evaluation layers. Intelligibility can be estimated by sending generated speech through an ASR model and measuring WER or CER. Speaker similarity estimates whether the voice identity is preserved. Human listening is still required for naturalness, rhythm, pronunciation, emotion, and long-form consistency.
An IndexTTS CV3-Eval table reports IndexTTS2.5 at 6.75% average WER and 73.18% speaker similarity, VoxCPM2 at 7.22% and 72.02%, and OmniVoice at 6.76% and 70.39%. These are vendor-reported project results under a particular evaluation setup and are not a universal ranking of Scribis custom TTS, Qwen3-TTS, Kokoro, or Piper.
Build a multilingual test script with English acronyms, Japanese names, Chinese proper nouns, numbers, punctuation, long sentences, and emotional instructions. Use the same text for every model. Then record pronunciation errors, unnatural pauses, speaker consistency, generation time, and the amount of manual editing required in the final video.
A practical Scribis workflow for global content
Use Scribis to turn a recording into an editable transcript or start from an existing script. Generate an English, Japanese, or Chinese voiceover with the integrated TTS options. Review the audio with the caption timeline, fix wording or pronunciation, and regenerate only the affected sections when the script changes.
If the main audience is global English-speaking viewers, begin with an English voice and evaluate clarity, pace, and regional neutrality. Add Japanese or Chinese versions when the distribution plan requires localization. If the content is for Taiwan, compare the Taiwan-accent option in Scribis with external voices using a local listening panel.
FAQ
What languages does Scribis custom TTS support?
The product capability list used for this article confirms Chinese, Japanese, and English, plus a Taiwan accent. The actual voice selection and availability should be checked in the current Scribis version.
Is Qwen3-TTS better than Scribis custom TTS?
They have different strengths. Qwen3-TTS publicly emphasizes voice cloning and voice design. Scribis custom TTS emphasizes supported languages, Taiwan accent availability, and integration with the transcript-to-caption-to-voiceover workflow. Compare them on the same scripts and final video use case.
Can Kokoro produce every English accent?
No. Accent availability depends on the selected voice and version. A model’s language label does not guarantee every regional accent.
Can I use Piper voices commercially?
Do not assume so. Review the current Piper repository and the specific voice model card or license before commercial deployment.
What is the best TTS for international videos?
Use the voice that produces natural pronunciation in the target language, fits the desired accent and style, and reduces production work. For many teams, workflow integration is as important as the model’s research score.
Conclusion: choose the voice and workflow that match the audience
Scribis custom TTS is the most direct starting point when your content needs English, Japanese, Chinese, and a Taiwan accent inside a transcription and video-production workflow. Qwen3-TTS adds multilingual voice design and cloning, Kokoro offers a lightweight fixed-voice route, and Piper remains a useful external local-deployment reference.
CTA: Create multilingual voiceovers from scripts and captions in Scribis, then compare English, Japanese, Chinese, and regional voice options in the context of your final content.