ASR Comparison

Whisper Large-v3 vs Breeze-ASR-25: Regional Speech and Code-Switching Compared

personJH LAI
calendar_today

Compare Whisper Large-v3 and Breeze-ASR-25 for regional speech, Mandarin-English code-switching, timestamps, and captions. Learn why regional benchmark design matters in Scribis.

Scope and product note: Whisper Large-v3 and Breeze-ASR-25 are integrated with Scribis. This article is written for an international audience, using Taiwanese Mandarin and Mandarin-English code-switching as a regional stress test rather than as a claim that every reader needs a Taiwan-specific model. The Twister research results shown below are not Breeze-ASR-25 results.

The short answer: Whisper is the global baseline; Breeze is a regional specialization

Whisper Large-v3 is a widely used multilingual speech-recognition baseline with an extensive ecosystem. Breeze-ASR-25 is a MediaTek Research model fine-tuned from Whisper-large-v2 with a focus on Traditional Chinese, Taiwanese Mandarin, Mandarin-English code-switching, and timestamp alignment.

For a global team, Whisper is usually the safer starting point when the input languages vary widely and the team values a mature ecosystem. Breeze-ASR-25 becomes interesting when a meaningful portion of the content contains Taiwanese Mandarin, Traditional Chinese terminology, or Mandarin-English switching. Both models are integrated with Scribis, so you can compare their outputs inside the same transcript and caption workflow.

Taiwanese Mandarin and mixed-language results from the Twister paper

Figure: Twister research results, not Breeze-ASR-25 results. MER combines Chinese CER and English WER in the cited paper.

What is Breeze-ASR-25?

Breeze-ASR-25 is a regional speech-to-text model designed for Chinese and English speech. Its model card highlights Traditional Chinese, Taiwanese Mandarin, and code-switching use cases, making it a useful specialist model for content teams serving Taiwan or working with Taiwanese speakers.

The practical distinction is not simply “Chinese versus English.” Regional speech includes pronunciation, vocabulary, numbers, names, borrowed English terms, and conversational patterns. A model can perform well on standard Mandarin and still require more corrections on Taiwanese Mandarin, local names, or mixed-language product discussions.

Breeze is therefore best framed as a targeted option. It is not a replacement for a global multilingual model in every market, but it may reduce editing effort for the regional content it was designed to serve.

Why Whisper Large-v3 remains a strong baseline

Whisper Large-v3 is often the first model teams test because it has broad language support, many deployment options, and a large collection of community tools. It is useful for multilingual interviews, lectures, podcasts, and caption pipelines where the language may not be known in advance.

The main weakness of a general baseline is that its average multilingual capability may hide regional gaps. If your recordings contain specialized accents or code-switching, evaluate those conditions directly. A regional model can be more useful even when it does not lead on a broad global benchmark.

Regional benchmark evidence: keep model and paper separate

The Twister paper evaluates a self-refining framework that improves Whisper-large-v2 with unlabelled speech, text, and synthetic speech. It is a research system and should not be presented as Breeze-ASR-25.

In the paper’s Table IV, the reported MER for CommonVoice16-zh-TW is 9.84 for Whisper-large-v2 auto, 8.95 for Whisper-large-v3, and 7.97 for Twister. On CSZS-zh-en, the corresponding results are 29.49, 26.43, and 13.01. On the ML-lecture-2021-long set, the reported values are 6.13, 6.41, and 4.98.

These numbers are useful because they show how code-switching can be substantially harder than clean regional speech. They do not prove that Breeze-ASR-25 has the same scores, because Breeze and Twister are different systems.

Dimension Whisper Large-v3 Breeze-ASR-25
Main positioning General multilingual baseline Regional Traditional Chinese and Taiwanese Mandarin specialist
Code-switching Broad capability; test the target language pair Explicit Mandarin-English focus in the model positioning
Ecosystem Very mature and broad More specialized, based on a Whisper fine-tuning route
Timestamp use Usually paired with alignment or wrapper tools Model positioning highlights timestamp alignment
Best use case Global multilingual content Regional Mandarin, Traditional Chinese, and Taiwan-oriented speech
Scribis status Integrated Integrated

How international teams should use this comparison

A global publisher does not need to choose one model for every country. You can use Whisper as a general fallback and route regional projects to a specialist model when the test set shows a meaningful reduction in corrections. This is especially useful for call centers, local news, education, field interviews, and product research.

For English-speaking readers, the lesson is broader than Taiwan. The same evaluation logic applies to Indian English, Scottish English, Singapore English, African English, Spanish-English code-switching, and any market where local speech differs from a textbook dataset. The right question is “which model handles our speakers and vocabulary?” rather than “which model has the largest language list?”

Caption alignment matters as much as recognition

A transcript can be correct while the captions still feel wrong. Caption quality depends on segment boundaries, reading speed, punctuation, speaker turns, and the placement of timestamps. When comparing Whisper and Breeze, track the number of manual caption cuts and merges, not only character or word errors.

Scribis is useful here because the model comparison can continue into timeline editing and export. A model that is slightly better on the transcript but produces poor segments may cost more editorial time than a model with similar recognition quality and better alignment.

A practical Scribis test plan

Prepare a small regional test set with clean speech, natural conversation, product names, numbers, English terms inside Mandarin sentences, and a longer lecture. Run both models on the same files. Compare recognition errors, timestamp drift, caption segmentation, and the number of edits required before publishing.

If your team publishes mostly global English and European-language content, Whisper may remain the simplest default. If a significant project segment contains Taiwanese Mandarin or Traditional Chinese, Breeze can be evaluated as a specialist route without changing the rest of the Scribis workflow.

FAQ

Is Breeze-ASR-25 better than Whisper Large-v3?

It can be a better fit for Taiwanese Mandarin and Traditional Chinese content, while Whisper is a broader general baseline. The answer depends on the speakers, vocabulary, language pair, and caption requirements.

Does the Twister chart show Breeze-ASR-25 performance?

No. The chart shows results from the Twister research framework and explicitly separates it from Breeze-ASR-25. It is included to illustrate why regional and code-switching benchmarks matter.

Can international teams use Breeze-ASR-25?

Yes, but it should be treated as a specialist model. Test it on the regional speech and language combinations that appear in your content rather than assuming that a Taiwan-focused model will improve every English recording.

Which model is better for captions?

Choose the model that reduces total caption editing time. That includes transcript corrections, segment repairs, timestamp adjustments, and export checks.

Conclusion: use a global baseline and add regional specialists where they help

Whisper Large-v3 is a strong general multilingual baseline. Breeze-ASR-25 is a targeted option for Traditional Chinese, Taiwanese Mandarin, and Mandarin-English code-switching. A global content operation can use both: Whisper for broad coverage and Breeze for regional projects where a specialist model reduces correction and alignment work.

CTA: Compare Whisper Large-v3 and Breeze-ASR-25 on your own regional or multilingual recordings in Scribis, then review the result as a finished caption workflow rather than a raw transcript alone.

References

  1. Breeze-ASR-25 official model card
  2. Twister research paper
  3. OpenAI Whisper official repository
  4. Scribis official website