arXiv:2412.16530cs.SDcs.CL2024-12中稿 · ICASSP, 4 pages被引 3

提升视频配音时口型与语音的同步性,让翻译后的人声更自然逼真。

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation

  • 在AVS2S模型训练中加入口型同步损失函数。
  • 跨四种语言对平均LSE-D降低9.2%,达10.67分。
  • 保持翻译质量与语音自然度,无性能下降。

音频-视觉语音到语音翻译通常侧重于提升翻译质量和语音自然度。然而,在音视频内容中,口型同步——即嘴唇动作与语音内容一致——对于保持配音视频的真实感至关重要。尽管如此,现有方法大多忽视了口型同步约束。本文通过在AVS2S模型训练中引入口型同步损失,填补这一空白。所提方法显著提升了直接音视频语音翻译中的口型同步性,四个语言对平均LSE-D得分降至10.67,相比强基线降低9.2%。同时,在原视频上叠加翻译语音时,仍保持高翻译质量与语音自然度,未造成任何性能退化。

原文摘要 · Abstract (English)

Audio-Visual Speech-to-Speech Translation typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony-ensuring that the movements of the lips match the spoken content-essential for maintaining realism in dubbed videos. Despite its importance, the inclusion of lip-synchrony constraints in AVS2S models has been largely overlooked. This study addresses this gap by integrating a lip-synchrony loss into the training process of AVS2S models. Our proposed method significantly enhances lip-synchrony in direct audio-visual speech-to-speech translation, achieving an average LSE-D score of 10.67, representing a 9.2% reduction in LSE-D over a strong baseline across four language pairs. Additionally, it maintains the naturalness and high quality of the translated speech when overlaid onto the original video, without any degradation in translation quality.

语音翻译口型同步多模态视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。