arXiv:2503.10488cs.LGcs.CV2025-03AAAI

实时生成语音同步手势,速度提升4倍且保持自然流畅。

Streaming Generation of Co-Speech Gestures via Accelerated Rolling Diffusion

  • 用分层噪声调度让扩散模型连续生成长序列动作。
  • 在两个真实数据集上均实现4倍加速,动作连贯性与真实感强。
  • 适合作为现有手势生成模型的通用加速插件,适合实时应用。

实时生成语音同步手势需兼顾时间一致性与采样效率。本文提出一种新型流式手势生成框架,通过结构化渐进噪声调度扩展滚动扩散模型,在保持动作真实性和多样性的同时,实现无缝长序列运动合成。该框架可兼容现有基于扩散的手势生成模型,将其转化为无需后处理的连续生成方法。我们在ZEGGS和BEAT两个强基准上评估,应用于顶尖基线模型时表现持续领先,证明其通用性与高效性。此外,我们提出滚动扩散梯级加速(RDLA)方法,采用梯级噪声调度策略同时去噪多帧,显著提升采样效率,实验中实现最高4倍加速,视觉保真度与时间连贯性俱佳。用户研究进一步验证了框架能生成与音频高度同步、真实且多样化的手势。

原文摘要 · Abstract (English)

Generating co-speech gestures in real time requires both temporal coherence and efficient sampling. We introduce a novel framework for streaming gesture generation that extends Rolling Diffusion models with structured progressive noise scheduling, enabling seamless long-sequence motion synthesis while preserving realism and diversity. Our framework is universally compatible with existing diffusion-based gesture generation model, transforming them into streaming methods capable of continuous generation without requiring post-processing. We evaluate our framework on ZEGGS and BEAT, strong benchmarks for real-world applicability. Applied to state-of-the-art baselines on both datasets, it consistently outperforms them, demonstrating its effectiveness as a generalizable and efficient solution for real-time co-speech gesture synthesis. We further propose Rolling Diffusion Ladder Acceleration (RDLA), a new approach that employs a ladder-based noise scheduling strategy to simultaneously denoise multiple frames. This significantly improves sampling efficiency while maintaining motion consistency, achieving up to a 4x speedup with high visual fidelity and temporal coherence in our experiments. Comprehensive user studies further validate our framework ability to generate realistic, diverse gestures closely synchronized with the audio input.

手势生成扩散模型实时生成加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。