Stop Syncing Subtitles Manually! 6 Updates to Save You from Tedious Editing Chores
To solve the workflow-interrupting hassle of subtitling, we've launched this wave of updates. No empty technical jargon, just 6 practical features to help you clock out earlier.
Stop Syncing Subtitles Manually! 6 Updates to Save You from Tedious Editing Chores
The most passion-draining part of video editing is often subtitling. After finally finishing recording a podcast or a multi-person interview, and throwing it into the editing software, you're faced with a string of speech-to-text transcripts where you can't tell who is who. You can only listen line by line, manually tag roles, and then fine-tune timelines that are off by just a few milliseconds.
To solve these workflow-interrupting hassles, we've launched this wave of updates. No empty technical jargon, just 6 practical features to help you clock out earlier:
Speech Recognition and Voice Processing
- Three Brand New ASR Models: We've added Paraformer-zh, Nemotron, and Qwen 3 ASR 1.7b. If you often deal with mixed Chinese and English terminology, or if the speaker mumbles or there's background noise, these models can drastically reduce the frequency of your post-production "manual typo correction".
- Voice Conversion Feature: If you feel the original recording is too dry, or if you want to add different narrator roles to your video but can't hire voice actors, this feature can directly convert existing speech into different voice qualities and intonations.
Subtitle Pacing and Visual Editing
- Characters Per Second (CPS) Segmentation: Traditional automatic subtitles often cut out sentences that are too long or too short, giving the audience no time to read. Now the system will automatically calculate based on Characters Per Second, segmenting subtitles at the most visually comfortable nodes.
- Visual Timeline: No more staring at numbers like "00:01:23.04" to manually input fine-tuning. We've visualized the time track. Whichever sentence you want to adjust, just drag the waveform edge with your mouse to align it, it's much more intuitive.
Multi-person Collaboration and Workflow
- Independent "Speaker" Management: When dealing with multi-person interviews, we help you extract the speakers independently and put them on a dedicated tab. You can see Speaker 1, Speaker 2 at a glance, and rename them uniformly with one click, without having to scroll up and down in hundreds of lines of transcript.
- Guided Workflow (Confirm Names First, Then Process Full Text): Before, when throwing the whole article to AI for translation or optimization, the biggest fear was that it "wrote Gary in the previous sentence and translated it to something else in the next". The new workflow takes a two-step approach: AI first finds and locks the "names" in the entire audio file, uses this roster as a baseline, and then translates or optimizes the full text. This completely solves the fatal flaw of character names conflicting with each other in long texts.
Next time you edit a video, try it like this:
First, throw the audio file into the workflow, let AI lock the names and optimize the transcript. Then go to the speaker tab to uniformly tag the roles, and turn on CPS segmentation to let the system automatically cut the subtitle pacing. Finally, fine-tune the waveform on the timeline, apply a voice you like, and call it a day.
Open the software update, and experience the feeling of not having to manually wrestle with the timeline.