SenseVoice Small vs Paraformer-zh: Which Fast ASR Model Is Best for Captions?
Compare SenseVoice Small and Paraformer-zh for fast ASR, captions, language identification, audio events, latency, and deployment. See how to test both in Scribis.
Scope and product note: SenseVoice Small and Paraformer-zh are integrated with Scribis. Fun-ASR-Nano is included as an external comparison reference, not as a Scribis integration. This article focuses on fast speech-to-text for captions, customer support, lectures, and multilingual media workflows.
The short answer: Paraformer is focused and lightweight; SenseVoice covers more speech tasks
SenseVoice Small and Paraformer-zh are both useful when speed matters. Paraformer-zh is an approximately 220M non-autoregressive ASR model focused on Chinese and English speech. SenseVoice Small is a broader speech foundation model that adds language identification, speech emotion recognition, and audio event detection to the ASR workflow.
If you need a fast caption draft, Paraformer-zh is a straightforward candidate. If you want to identify languages or annotate additional audio events, SenseVoice Small has a wider task scope. Both are integrated with Scribis, where the result can move from transcription into caption timing, translation, editing, and export.

Figure: Qualitative feature positioning, not a numerical ranking. Fun-ASR-Nano is labelled as an external comparison model.
What is SenseVoice Small?
SenseVoice Small is part of the FunASR ecosystem. Its official model card describes a non-autoregressive end-to-end architecture trained on more than 400,000 hours of data and positioned for low-latency multilingual speech understanding.
Its most important distinction is that it is not only an ASR checkpoint. Depending on the pipeline, SenseVoice can provide speech transcription, language identification, emotional attributes, and audio-event information. This can be useful for media indexing, moderation support, meeting analysis, and audio search.
FunASR documentation has reported approximately 70ms inference time for 10 seconds of audio in a particular setup and described it as much faster than Whisper Large. That is a vendor-reported test result; hardware, batch size, audio format, and runtime can change the outcome.
What is Paraformer-zh?
Paraformer-zh uses a SANM encoder, a CIF predictor, and a non-autoregressive decoder. Its design emphasizes fast, predictable transcription rather than a large set of secondary speech-understanding tasks.
That narrower scope is often an advantage. A caption pipeline may not need emotion labels or audio-event tags. A smaller and more focused model can be easier to run at scale, easier to quantize, and easier to understand when a team is troubleshooting recognition errors.
| Dimension | SenseVoice Small | Paraformer-zh |
|---|---|---|
| Main positioning | Multitask speech foundation model | Lightweight Chinese/English ASR |
| ASR | Supported | Supported |
| Language identification | A core task in the model positioning | Usually handled by an additional pipeline or detector |
| Emotion and audio events | Supported in the documented ecosystem | Not the primary focus |
| Architecture | Non-autoregressive end-to-end | SANM + CIF + non-autoregressive decoder |
| Best use case | Fast transcription plus speech analytics | Captions, customer support, batch, and edge deployment |
| Scribis status | Integrated | Integrated |
Comparison with Fun-ASR-Nano, Parakeet, and Qwen3-ASR
Fun-ASR-Nano-2512 is an external model in this article. It is designed for Chinese dialects, English, Japanese, and industry speech, making it a useful comparison when the content is multilingual or domain-specific. Its broader coverage may be valuable, but it should not be described as a Scribis integration.
Parakeet TDT v3 represents another efficiency-oriented path. A third-party English short-form leaderboard reports Parakeet TDT v3 at 6.32% average WER and 3,330 RTFx on a fixed A100 setup. That does not mean it is the best model for Chinese captions; it shows why throughput, language, and caption quality need separate tests.
Qwen3-ASR adds 30 languages and 22 Chinese dialects or accents, with offline and streaming workflows. It is a useful integrated comparison point when a fast Chinese pipeline must also handle multilingual content or regional speech.
| Model | Best fit | Difference from SenseVoice and Paraformer |
|---|---|---|
| SenseVoice Small | ASR plus language and audio analysis | Broader task coverage |
| Paraformer-zh | Fast Chinese captions and customer support | More focused and lightweight |
| Fun-ASR-Nano | Chinese dialects, English, Japanese, and domain speech | External multilingual comparison model |
| Parakeet TDT v3 | High-throughput batch processing | Strong speed reference; test target language separately |
| Qwen3-ASR | Multilingual and dialect-aware transcription | More language and streaming emphasis |
Why caption quality needs more than WER
A transcript may have an acceptable character or word error rate and still produce poor captions. Captioning also depends on sentence boundaries, reading speed, punctuation, speaker turns, and timestamp behavior. A fast model that creates badly segmented lines can increase editorial cost.
Test at least a clean recording, a noisy mobile recording, a two-person conversation, a lecture, and a file with names, numbers, and product terms. Measure recognition errors, caption segmentation fixes, timing adjustments, and the number of minutes a human editor needs before export.
If SenseVoice’s additional language or audio-event information is not useful to your project, Paraformer-zh may offer a simpler pipeline. If those signals help search or analysis, SenseVoice may create more value even when the ASR text is similar.
A practical Scribis workflow
Import MP3, WAV, M4A, or MP4 content into Scribis and create a first transcript with SenseVoice Small or Paraformer-zh. Review the output in the caption timeline, then check punctuation, segment length, translation, speaker labels, and export format.
A useful operating model is to use SenseVoice for rapid triage and audio analysis, then use Paraformer-zh or Qwen3-ASR on files that require more focused caption quality. Since both primary models are integrated with Scribis, the team can compare the output in one content workflow rather than maintaining separate manual pipelines.
FAQ
Is SenseVoice Small more accurate than Paraformer-zh?
There is no universal answer. SenseVoice focuses on multitask speech understanding and low latency, while Paraformer-zh is a focused Chinese/English ASR model. Accuracy depends on language, noise, domain, and post-processing.
Is Fun-ASR-Nano integrated with Scribis?
Not according to the current integration list used for this content set. Fun-ASR-Nano is an external comparison reference in this article.
Can SenseVoice detect emotions reliably?
It provides documented emotion and audio-event capabilities, but automated labels should be treated as model outputs rather than ground truth. Validate them for your use case.
Which model is better for fast captions?
Use Paraformer-zh as a focused fast-caption candidate and SenseVoice Small when additional speech analysis matters. The final decision should use caption correction time and export quality, not only inference speed.
Conclusion: choose a focused ASR or a wider speech pipeline
Paraformer-zh is a strong candidate for lightweight, fast Chinese and English captions. SenseVoice Small is more attractive when transcription must be combined with language identification, emotion, or audio-event analysis. Fun-ASR-Nano, Parakeet, and Qwen3-ASR provide useful external or integrated comparison points for dialects, throughput, and multilingual coverage.
CTA: Compare SenseVoice Small and Paraformer-zh on the same files in Scribis, then evaluate the finished captions, timeline, and export—not just the raw transcript.