用眼睛鼻子嘴控制面部动作,让静态人像动起来更自然。
Stable Video-Driven Portraits
- 用驱动视频中眼部口鼻区域作为强运动信号,精准控制表情变化。
- 生成视频时长稳定、表情连贯,避免身份泄露和画面抖动。
- 适合影视特效、虚拟主播等需要高保真人脸动画的场景。
人脸动画旨在从一张源图像生成逼真视频,通过复现驱动视频中的表情与姿态。早期方法依赖3D可变形模型或特征扭曲技术,常面临表现力有限、时间不一致及对未知身份或大幅姿态变化泛化能力差的问题。近期基于扩散模型的方法虽提升质量,但仍受限于控制信号弱和架构瓶颈。本文提出一种新型扩散框架,利用驱动视频中眼部、鼻部和口部的掩码区域作为强运动控制信号。为防止外观信息泄露,采用跨身份监督训练。为充分利用预训练扩散模型的先验知识,模型引入极少新参数,收敛更快且泛化性更好。设计时空注意力机制,实现帧间与帧内交互,有效捕捉细微动作并减少时间伪影。推理时引入新颖信号融合策略,平衡动作保真度与身份一致性。所提方法在时间连续性和表达控制上表现优异,适用于真实场景的高质量可控人脸动画。
原文摘要 · Abstract (English)
Portrait animation aims to generate photo-realistic videos from a single source image by reenacting the expression and pose from a driving video. While early methods relied on 3D morphable models or feature warping techniques, they often suffered from limited expressivity, temporal inconsistency, and poor generalization to unseen identities or large pose variations. Recent advances using diffusion models have demonstrated improved quality but remain constrained by weak control signals and architectural limitations. In this work, we propose a novel diffusion based framework that leverages masked facial regions specifically the eyes, nose, and mouth from the driving video as strong motion control cues. To enable robust training without appearance leakage, we adopt cross identity supervision. To leverage the strong prior from the pretrained diffusion model, our novel architecture introduces minimal new parameters that converge faster and help in better generalization. We introduce spatial temporal attention mechanisms that allow inter frame and intra frame interactions, effectively capturing subtle motions and reducing temporal artifacts. Our model uses history frames to ensure continuity across segments. At inference, we propose a novel signal fusion strategy that balances motion fidelity with identity preservation. Our approach achieves superior temporal consistency and accurate expression control, enabling high-quality, controllable portrait animation suitable for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。