arXiv:2512.17331cs.CV2025-12

用注意力协同变形,让虚拟人像口型更自然、细节更完整。

SynergyWarpNet: Attention-Guided Cooperative Warping for Neural Portrait Animation

  • 分三步:先粗对齐,再用参考图补全遮挡区,最后自适应融合结果
  • 在多个基准数据集上达到当前最好效果,口型同步更精准
  • 适合做虚拟主播、数字人动画的开发者参考

神经肖像动画在虚拟形象、远程通信和数字内容创作中展现出巨大潜力。然而,传统显式变形方法常难以准确传递动作或恢复缺失区域,而现有基于注意力的变形方法虽有效,却往往计算复杂且缺乏几何约束。为此,我们提出SynergyWarpNet——一种注意力引导的协同变形框架,用于高保真说话头像生成。给定源肖像、驱动图像及一组参考图像,模型分三阶段逐步优化动画:首先,通过3D密集光流实现源图与驱动图的粗略空间对齐;其次,参考增强校正模块利用多参考图像的3D关键点与纹理特征间的交叉注意力,语义填补遮挡或扭曲区域;最后,置信度引导融合模块采用空间自适应融合策略,结合学习到的置信度图平衡结构对齐与视觉一致性。在多个基准数据集上的全面评估表明,该方法性能达到当前最优水平。

原文摘要 · Abstract (English)

Recent advances in neural portrait animation have demonstrated remarked potential for applications in virtual avatars, telepresence, and digital content creation. However, traditional explicit warping approaches often struggle with accurate motion transfer or recovering missing regions, while recent attention-based warping methods, though effective, frequently suffer from high complexity and weak geometric grounding. To address these issues, we propose SynergyWarpNet, an attention-guided cooperative warping framework designed for high-fidelity talking head synthesis. Given a source portrait, a driving image, and a set of reference images, our model progressively refines the animation in three stages. First, an explicit warping module performs coarse spatial alignment between the source and driving image using 3D dense optical flow. Next, a reference-augmented correction module leverages cross-attention across 3D keypoints and texture features from multiple reference images to semantically complete occluded or distorted regions. Finally, a confidence-guided fusion module integrates the warped outputs with spatially-adaptive fusing, using a learned confidence map to balance structural alignment and visual consistency. Comprehensive evaluations on benchmark datasets demonstrate state-of-the-art performance.

人脸动画注意力机制图像变形虚拟人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。