用扩散模型提升2D说话人换脸的肢体与纹理细节控制
TALK-Act: Enhance Textural-Awareness for 2D Speaking Avatar Reenactment with Diffusion Model
- 通过运动引导增强扩散模型的纹理感知能力
- 仅需30秒视频即可实现高保真换脸重演
- 适合需要精细肢体动作控制的虚拟形象应用
近年来,随着人脸动画技术的快速发展,2D说话人已广泛应用于日常场景。然而,现有方法大多忽略对人物身体的显式控制。本文提出TALK-Act框架,不仅驱动面部,还同步控制躯干和手势运动。受扩散模型启发,我们设计了运动增强的纹理感知建模机制,通过构建2D与3D结构信息作为中间引导。不同于以往依赖侧网络注入控制信息的方法,我们的方法在人物特定微调后仍能生成时序稳定的图像。引入运动增强的纹理对齐模块以强化驱动信号与目标信号间的关联,并构建基于记忆的手部恢复模块以解决手形保持难题。预训练后,仅需30秒个人视频数据即可实现高保真2D人物重演。大量实验证明本框架的有效性与优越性。
原文摘要 · Abstract (English)
Recently, 2D speaking avatars have increasingly participated in everyday scenarios due to the fast development of facial animation techniques. However, most existing works neglect the explicit control of human bodies. In this paper, we propose to drive not only the faces but also the torso and gesture movements of a speaking figure. Inspired by recent advances in diffusion models, we propose the Motion-Enhanced Textural-Aware ModeLing for SpeaKing Avatar Reenactment (TALK-Act) framework, which enables high-fidelity avatar reenactment from only short footage of monocular video. Our key idea is to enhance the textural awareness with explicit motion guidance in diffusion modeling. Specifically, we carefully construct 2D and 3D structural information as intermediate guidance. While recent diffusion models adopt a side network for control information injection, they fail to synthesize temporally stable results even with person-specific fine-tuning. We propose a Motion-Enhanced Textural Alignment module to enhance the bond between driving and target signals. Moreover, we build a Memory-based Hand-Recovering module to help with the difficulties in hand-shape preserving. After pre-training, our model can achieve high-fidelity 2D avatar reenactment with only 30 seconds of person-specific data. Extensive experiments demonstrate the effectiveness and superiority of our proposed framework. Resources can be found at https://guanjz20.github.io/projects/TALK-Act.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。