用细粒度偏好优化解决多语言配音时长不匹配问题
Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
- 将配音时长对齐建模为分段偏好优化问题
- 在多个数据集上显著提升语音与画面同步效果
- 适合需要精准音画同步的影视翻译场景
视频配音旨在将视觉媒体中的源语言语音翻译为目标语言,依赖神经机器翻译和文本转语音技术。由于语言间信息密度差异,目标语音时常与源语音时长不匹配,导致音画不同步,严重影响观看体验。本研究将基于大模型的视频配音时长对齐问题视为偏好优化任务,提出分段监督偏好优化(SSPO)方法,采用分段采样策略与细粒度损失函数,有效缓解源语句与目标语句间的时长偏差。实验结果表明,SSPO在时长对齐任务中表现更优。
原文摘要 · Abstract (English)
Video dubbing aims to translate original speech in visual media programs from the source language to the target language, relying on neural machine translation and text-to-speech technologies. Due to varying information densities across languages, target speech often mismatches the source speech duration, causing audio-video synchronization issues that significantly impact viewer experience. In this study, we approach duration alignment in LLM-based video dubbing machine translation as a preference optimization problem. We propose the Segment Supervised Preference Optimization (SSPO) method, which employs a segment-wise sampling strategy and fine-grained loss to mitigate duration mismatches between source and target lines. Experimental results demonstrate that SSPO achieves superior performance in duration alignment tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。