arXiv:2501.04586cs.CV2025-01IJCV被引 2

用动作变形技术实现配音时口型同步且保留原人物特征

Identity-Preserving Video Dubbing Using Motion Warping

  • 通过变换器对齐音频与图像,动态捕捉音视频对应关系
  • 采用运动扭曲策略精准匹配目标口型,提升生成质量
  • 适合需要高保真人物特征保留的影视配音场景

视频配音旨在从参考视频和驱动音频中合成逼真的唇同步视频。现有方法虽能准确生成由音频驱动的嘴型,但常无法保持身份特异性特征,因未能有效捕捉音频线索与参考身份视觉属性间的细微关联。导致生成结果在再现参考身份的独特纹理和结构细节方面缺乏真实感。为此,我们提出IPTalker,一种新颖且鲁棒的视频配音框架,在确保唇同步准确性的同时,实现驱动音频与参考身份的无缝对齐。核心是基于变换器的对齐机制,动态建模音频特征与参考图像间的对应关系,实现精确的身份感知音视频融合。在此基础上,运动扭曲策略进一步通过空间形变参考图像以匹配目标音频驱动配置。专门设计的优化流程则缓解遮挡伪影,增强细粒度纹理(如嘴部细节、皮肤特征)的保留。大量定性和定量评估表明,IPTalker在真实性、唇同步和身份保留方面持续优于现有方法,确立了高质量、身份一致视频配音的新基准。

原文摘要 · Abstract (English)

Video dubbing aims to synthesize realistic, lip-synced videos from a reference video and a driving audio signal. Although existing methods can accurately generate mouth shapes driven by audio, they often fail to preserve identity-specific features, largely because they do not effectively capture the nuanced interplay between audio cues and the visual attributes of reference identity . As a result, the generated outputs frequently lack fidelity in reproducing the unique textural and structural details of the reference identity. To address these limitations, we propose IPTalker, a novel and robust framework for video dubbing that achieves seamless alignment between driving audio and reference identity while ensuring both lip-sync accuracy and high-fidelity identity preservation. At the core of IPTalker is a transformer-based alignment mechanism designed to dynamically capture and model the correspondence between audio features and reference images, thereby enabling precise, identity-aware audio-visual integration. Building on this alignment, a motion warping strategy further refines the results by spatially deforming reference images to match the target audio-driven configuration. A dedicated refinement process then mitigates occlusion artifacts and enhances the preservation of fine-grained textures, such as mouth details and skin features. Extensive qualitative and quantitative evaluations demonstrate that IPTalker consistently outperforms existing approaches in terms of realism, lip synchronization, and identity retention, establishing a new state of the art for high-quality, identity-consistent video dubbing.

视频配音口型同步身份保留运动扭曲

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。