用内容感知编辑提升配音视频的唇形同步与画面自然度
Video Editing for Audio-Visual Dubbing
- 将配音任务重构为内容感知编辑,保留原视频上下文
- 在遮挡唇部数据集上实现更高身份保真度与同步精度
- 适合需要高质量跨语言视频重制的创作者
视觉配音(visual dubbing)通过使面部动作与新语音同步,对实现多语言内容无障碍传播至关重要。现有方法常生成孤立的说话人脸,难以融入原场景,或使用修补技术丢失关键视觉信息如部分遮挡和光照变化。本文提出EdiDub框架,将视觉配音重新定义为内容感知编辑任务。该框架通过专用条件机制,在不复制原图的前提下忠实还原原始视频上下文。在多个基准测试中,包括具有挑战性的遮挡唇部数据集,EdiDub显著提升了身份保持率和同步性能。人工评估进一步验证其优势,相比领先方法在同步性和视觉自然度评分上均更优。结果表明,内容感知编辑优于传统生成或修补方法,尤其在保持复杂视觉元素的同时确保精准唇形同步。
原文摘要 · Abstract (English)
Visual dubbing, the synchronization of facial movements with new speech, is crucial for making content accessible across different languages, enabling broader global reach. However, current methods face significant limitations. Existing approaches often generate talking faces, hindering seamless integration into original scenes, or employ inpainting techniques that discard vital visual information like partial occlusions and lighting variations. This work introduces EdiDub, a novel framework that reformulates visual dubbing as a content-aware editing task. EdiDub preserves the original video context by utilizing a specialized conditioning scheme to ensure faithful and accurate modifications rather than mere copying. On multiple benchmarks, including a challenging occluded-lip dataset, EdiDub significantly improves identity preservation and synchronization. Human evaluations further confirm its superiority, achieving higher synchronization and visual naturalness scores compared to the leading methods. These results demonstrate that our content-aware editing approach outperforms traditional generation or inpainting, particularly in maintaining complex visual elements while ensuring accurate lip synchronization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。