让配音音频精准匹配原声时长,支持实时本地化处理
Length Aware Speech Translation for Video Dubbing
- 用音素级模型+长度标签实现长短句统一翻译
- 单次解码生成不同长度译文,同步质量提升显著
- 适合移动端实时视频配音,尤其对韩语提升明显
在视频配音中,使翻译后的音频与源音频对齐是一项重大挑战。本文针对实时、本地化视频配音场景,提出一种基于音素的端到端长度感知语音翻译(LSST)模型,通过预定义标签生成短、正常、长三种长度的翻译。同时引入长度感知束搜索(LABS),可在一次解码过程中高效生成不同长度的译文。该方法在保持与无长度感知基线相当的BLEU分数的同时,显著提升了源音频与目标音频之间的同步质量,西班牙语和韩语的平均意见得分(MOS)分别提升0.34和0.65。
原文摘要 · Abstract (English)
In video dubbing, aligning translated audio with the source audio is a significant challenge. Our focus is on achieving this efficiently, tailored for real-time, on-device video dubbing scenarios. We developed a phoneme-based end-to-end length-sensitive speech translation (LSST) model, which generates translations of varying lengths short, normal, and long using predefined tags. Additionally, we introduced length-aware beam search (LABS), an efficient approach to generate translations of different lengths in a single decoding pass. This approach maintained comparable BLEU scores compared to a baseline without length awareness while significantly enhancing synchronization quality between source and target audio, achieving a mean opinion score (MOS) gain of 0.34 for Spanish and 0.65 for Korean, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。