arXiv:2508.20210cs.CV2025-08被引 10

解决音频驱动人体动画长期生成中的失真与手部不准问题

InfinityHuman: Towards Long-Term Audio-Driven Human

  • 分阶段生成:先同步音频特征,再用姿态引导细化视频
  • 在EMTD和HDTF数据集上实现最佳画质与口型同步效果
  • 引入手部专属奖励机制,提升手势真实感,适合影视动画应用

音频驱动人体动画因实用价值广受关注,但生成高分辨率、长时长视频仍面临外观不一致与手部动作不自然的挑战。现有方法通过重叠运动帧扩展视频,但存在误差累积,导致身份漂移、色彩偏移和场景不稳定;且手部运动建模不足,出现明显扭曲与音频不同步。本文提出InfinityHuman,一种从粗到精的框架:先生成音画同步表示,再通过姿态引导的精修模块逐步生成高清长视频。由于姿态序列与外观解耦且抗时间退化,该模块以稳定姿态和初始帧为视觉锚点,有效减少漂移并提升口型同步。此外,引入基于高质量手部动作数据训练的手部专属奖励机制,增强语义准确性和手势自然度。在EMTD和HDTF数据集上的实验表明,InfinityHuman在视频质量、身份保持、手部精度和口型同步方面均达到当前最优表现。消融实验进一步验证各模块有效性。代码将公开。

原文摘要 · Abstract (English)

Audio-driven human animation has attracted wide attention thanks to its practical applications. However, critical challenges remain in generating high-resolution, long-duration videos with consistent appearance and natural hand motions. Existing methods extend videos using overlapping motion frames but suffer from error accumulation, leading to identity drift, color shifts, and scene instability. Additionally, hand movements are poorly modeled, resulting in noticeable distortions and misalignment with the audio. In this work, we propose InfinityHuman, a coarse-to-fine framework that first generates audio-synchronized representations, then progressively refines them into high-resolution, long-duration videos using a pose-guided refiner. Since pose sequences are decoupled from appearance and resist temporal degradation, our pose-guided refiner employs stable poses and the initial frame as a visual anchor to reduce drift and improve lip synchronization. Moreover, to enhance semantic accuracy and gesture realism, we introduce a hand-specific reward mechanism trained with high-quality hand motion data. Experiments on the EMTD and HDTF datasets show that InfinityHuman achieves state-of-the-art performance in video quality, identity preservation, hand accuracy, and lip-sync. Ablation studies further confirm the effectiveness of each module. Code will be made public.

音频驱动人体动画手部生成长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。