arXiv:2506.16852cs.CV2025-06ICCV被引 2

用扩散模型实现可控制的表情头像一键替换,支持换脸后自由调整表情。

Controllable and Expressive One-Shot Video Head Swapping

  • 基于统一扩散框架,分离头部身份与背景,保持整体头型一致。
  • 能精准迁移表情,支持真实与虚拟角色的高质量表达传递。
  • 适合影视特效、虚拟形象定制等需要精细控制头像表达的场景。

本文提出一种基于扩散模型的多条件可控视频头像替换框架,可将静态图像中的人脸无缝移植到动态视频中,保留目标视频的原始身体和背景,并支持在替换过程中灵活调整头部表情与动作。现有方法大多只关注局部面部替换,忽略整体头部形态;而现有头像替换技术在发式多样性和复杂背景下表现不佳,且无法在替换后修改表情。为此,我们提出创新策略:1)身份保真上下文融合:采用与形状无关的掩码策略,显式分离前景头部身份特征与背景/身体上下文,并结合发型增强策略,实现对多种发式和复杂背景下头部身份的鲁棒保留;2)表达感知的地标重映射与编辑:提出解耦的3DMM驱动重映射模块,分离身份、表情与头部姿态,降低输入图像原始表情的影响,支持表情编辑;同时采用尺度感知重映射策略,减少跨身份表情失真,提升转移精度。实验表明,该方法在背景融合自然性、源肖像身份保留以及表情迁移能力方面均表现优异,适用于真实与虚拟角色。

原文摘要 · Abstract (English)

In this paper, we propose a novel diffusion-based multi-condition controllable framework for video head swapping, which seamlessly transplant a human head from a static image into a dynamic video, while preserving the original body and background of target video, and further allowing to tweak head expressions and movements during swapping as needed. Existing face-swapping methods mainly focus on localized facial replacement neglecting holistic head morphology, while head-swapping approaches struggling with hairstyle diversity and complex backgrounds, and none of these methods allow users to modify the transplanted head expressions after swapping. To tackle these challenges, our method incorporates several innovative strategies through a unified latent diffusion paradigm. 1) Identity-preserving context fusion: We propose a shape-agnostic mask strategy to explicitly disentangle foreground head identity features from background/body contexts, combining hair enhancement strategy to achieve robust holistic head identity preservation across diverse hair types and complex backgrounds. 2) Expression-aware landmark retargeting and editing: We propose a disentangled 3DMM-driven retargeting module that decouples identity, expression, and head poses, minimizing the impact of original expressions in input images and supporting expression editing. While a scale-aware retargeting strategy is further employed to minimize cross-identity expression distortion for higher transfer precision. Experimental results demonstrate that our method excels in seamless background integration while preserving the identity of the source portrait, as well as showcasing superior expression transfer capabilities applicable to both real and virtual characters.

视频生成头像替换扩散模型表情控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。