用检索增强扩散模型,让语音驱动的手势更自然、有表现力。
ExGes: Expressive Human Motion Retrieval and Modulation for Audio-Driven Gesture Synthesis
- 构建动作库并用对比学习精准检索参考姿态
- 在BEAT2数据集上降低6.2%的姿势距离,提升5.3%多样性
- 适合虚拟形象、人机交互等需要细腻手势的应用
语音驱动的人体手势生成在虚拟形象、人机交互和内容创作中具有广泛应用。尽管已有进展,现有方法常生成粗糙、缺乏表现力且与音频语义对不齐的手势。为此,我们提出ExGes,一种基于检索增强的扩散框架,包含三个关键设计:(1) 动作库构建,利用训练数据集建立手势库;(2) 动作检索模块,采用对比学习与动量蒸馏实现细粒度参考姿态检索;(3) 精确控制模块,结合部分掩码与随机掩码,实现灵活精细的控制。在BEAT2数据集上的实验表明,ExGes相比EMAGE将弗雷歇手势距离降低6.2%,运动多样性提升5.3%。用户研究显示,71.3%的参与者更偏好其自然性与语义相关性。代码将在录用后发布。
原文摘要 · Abstract (English)
Audio-driven human gesture synthesis is a crucial task with broad applications in virtual avatars, human-computer interaction, and creative content generation. Despite notable progress, existing methods often produce gestures that are coarse, lack expressiveness, and fail to fully align with audio semantics. To address these challenges, we propose ExGes, a novel retrieval-enhanced diffusion framework with three key designs: (1) a Motion Base Construction, which builds a gesture library using training dataset; (2) a Motion Retrieval Module, employing constrative learning and momentum distillation for fine-grained reference poses retreiving; and (3) a Precision Control Module, integrating partial masking and stochastic masking to enable flexible and fine-grained control. Experimental evaluations on BEAT2 demonstrate that ExGes reduces Fréchet Gesture Distance by 6.2\% and improves motion diversity by 5.3\% over EMAGE, with user studies revealing a 71.3\% preference for its naturalness and semantic relevance. Code will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。