用音频和手势驱动生成逼真半身人物动画,解决手部模糊和头部僵硬问题。
VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
- 预训练手部清晰码本,提升手部细节真实度
- 分离建模口型与头部动作,实现自然同步
- 姿势校准技巧让动作过渡更流畅,适合影视动画应用
现有音频与姿态驱动的人体动画方法常因音频与头部动作关联弱、手部结构复杂导致头部动作僵硬、手部模糊。为此,我们提出VividAnimator,一个端到端的半身人体动画生成框架,通过三种创新改进:首先,预训练手部清晰码本(HCC),编码高保真手部纹理先验,显著缓解手部退化;其次,设计双流音频感知模块(DSAA),分别建模唇同步与自然头部姿态动态,同时支持交互;第三,引入姿势校准技巧(PCT),通过放松刚性约束,精炼并对齐姿态条件,确保手势过渡平滑自然。大量实验表明,VividAnimator在定量指标与定性评估中均达到当前最优表现,生成视频具备卓越的手部细节、动作真实性与身份一致性。
原文摘要 · Abstract (English)
Existing for audio- and pose-driven human animation methods often struggle with stiff head movements and blurry hands, primarily due to the weak correlation between audio and head movements and the structural complexity of hands. To address these issues, we propose VividAnimator, an end-to-end framework for generating high-quality, half-body human animations driven by audio and sparse hand pose conditions. Our framework introduces three key innovations. First, to overcome the instability and high cost of online codebook training, we pre-train a Hand Clarity Codebook (HCC) that encodes rich, high-fidelity hand texture priors, significantly mitigating hand degradation. Second, we design a Dual-Stream Audio-Aware Module (DSAA) to model lip synchronization and natural head pose dynamics separately while enabling interaction. Third, we introduce a Pose Calibration Trick (PCT) that refines and aligns pose conditions by relaxing rigid constraints, ensuring smooth and natural gesture transitions. Extensive experiments demonstrate that Vivid Animator achieves state-of-the-art performance, producing videos with superior hand detail, gesture realism, and identity consistency, validated by both quantitative metrics and qualitative evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。