用反馈机制实现无需训练的长时间人脸手势动画生成
TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model
- 基于图像扩散模型构建反馈驱动机制,提升动作连贯性
- 可生成无限时长的上半身动画,无额外计算开销
- 适合需要长期一致动作生成的影视或虚拟角色应用
扩散模型的进展显著提升了角色驱动动画的真实感与泛化能力,仅需单张RGB图像和驱动姿态即可合成高质量动作。然而,生成长时间、时间连贯的内容仍具挑战。现有方法受限于计算与内存,通常在短视频片段上训练,仅能有效处理有限帧数,难以支持持续生成。为此,我们提出TalkingPose,一种专为生成长时间、时间一致的人类上半身动画而设计的新型扩散框架。该框架利用驱动帧精准捕捉面部与手部表情动作,并通过稳定的扩散主干网络无缝传递至目标演员。为确保动作连续并增强时间一致性,我们引入基于图像扩散模型的反馈驱动机制。该机制不增加额外计算成本,也无需二次训练,可实现无限时长动画生成。此外,我们构建了一个大规模、全面的上半身动画数据集,作为新基准。
原文摘要 · Abstract (English)
Recent advancements in diffusion models have significantly improved the realism and generalizability of character-driven animation, enabling the synthesis of high-quality motion from just a single RGB image and a set of driving poses. Nevertheless, generating temporally coherent long-form content remains challenging. Existing approaches are constrained by computational and memory limitations, as they are typically trained on short video segments, thus performing effectively only over limited frame lengths and hindering their potential for extended coherent generation. To address these constraints, we propose TalkingPose, a novel diffusion-based framework specifically designed for producing long-form, temporally consistent human upper-body animations. TalkingPose leverages driving frames to precisely capture expressive facial and hand movements, transferring these seamlessly to a target actor through a stable diffusion backbone. To ensure continuous motion and enhance temporal coherence, we introduce a feedback-driven mechanism built upon image-based diffusion models. Notably, this mechanism does not incur additional computational costs or require secondary training stages, enabling the generation of animations with unlimited duration. Additionally, we introduce a comprehensive, large-scale dataset to serve as a new benchmark for human upper-body animation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。