arXiv:2503.22225cs.CV2025-03被引 1

通过轨迹引导保持人脸编辑的时序一致性,让虚拟人物说话更自然。

Follow Your Motion: A Generic Temporal Consistency Portrait Editing Framework with Trajectory Guidance

  • 基于3D高斯泼溅渲染帧,用扩散模型学习多尺度运动轨迹。
  • 动态加权注意力机制使表情变化更连贯,提升细微动作一致性。
  • 适用于文本驱动、光影调整等场景,可修复已有不一致结果。

预训练的条件扩散模型在图像编辑中表现出巨大潜力,但在说话头场景中常面临时序不一致问题,因面部表情持续变化而加剧。这主要源于单帧独立编辑及编辑过程中的时序连续性丢失。本文提出通用框架Follow Your Motion(FYM),针对预训练3D高斯泼溅模型生成的人脸图像,首先构建扩散模型,从首帧到各后续帧直观且内在地学习不同尺度与像素坐标上的运动轨迹变化,确保编辑后的动态角色继承原始渲染序列的运动信息。其次,为实现说话头编辑中精细表情的时序一致性,提出动态重加权注意力机制,根据关键点损失动态调整空间关键点权重,从而获得更一致且精细的面部表情。大量实验证明,本方法在时序一致性方面优于现有方法,可广泛用于文本驱动编辑、光照调整及其他应用,并能优化和补偿多种场景中的时序不一致输出。

原文摘要 · Abstract (English)

Pre-trained conditional diffusion models have demonstrated remarkable potential in image editing. However, they often face challenges with temporal consistency, particularly in the talking head domain, where continuous changes in facial expressions intensify the level of difficulty. These issues stem from the independent editing of individual images and the inherent loss of temporal continuity during the editing process. In this paper, we introduce Follow Your Motion (FYM), a generic framework for maintaining temporal consistency in portrait editing. Specifically, given portrait images rendered by a pre-trained 3D Gaussian Splatting model, we first develop a diffusion model that intuitively and inherently learns motion trajectory changes at different scales and pixel coordinates, from the first frame to each subsequent frame. This approach ensures that temporally inconsistent edited avatars inherit the motion information from the rendered avatars. Secondly, to maintain fine-grained expression temporal consistency in talking head editing, we propose a dynamic re-weighted attention mechanism. This mechanism assigns higher weight coefficients to landmark points in space and dynamically updates these weights based on landmark loss, achieving more consistent and refined facial expressions. Extensive experiments demonstrate that our method outperforms existing approaches in terms of temporal consistency and can be used to optimize and compensate for temporally inconsistent outputs in a range of applications, such as text-driven editing, relighting, and various other applications.

人脸编辑时序一致性扩散模型3D高斯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。