arXiv:2608.19055cs.CVcs.GR2026-08

用扩散模型生成高精度鼓手动作,音频驱动且真实自然。

Generalized Audio-Driven Synthesis of Precise Drummer Motion

论文配图:Generalized Audio-Driven Synthesis of Precise Drummer Motion
图 1 · 摘自论文原文
  • 用双目标损失解耦身体动态与鼓棒精度,实现厘米级精准控制。
  • 在非剪辑的真实音频上仍能泛化,生成动作与真人表演难以区分。
  • 提出新评估指标,量化空间精度与音画同步性,填补领域空白。

音乐驱动的角色动画在娱乐和交互教育中具有变革潜力,但基于音频生成逼真打鼓动作仍具挑战,源于高加速度动态与极致时空精度之间的矛盾。现有方法多依赖动作匹配或MIDI输入,难以泛化至真实世界音频。同时,该领域缺乏能区分精确打鼓与噪声动作的标准评估指标。本文提出一种生成式扩散框架,采用双目标损失函数,解耦骨骼完整性与鼓棒精度,实现厘米级鼓棒定位,同时保持自然的身体动态。通过自建数据集与数据增强策略,模型可泛化至非剪辑的“野生”音频。为严格评估性能,我们提出两项新指标:冲击点到目标距离(量化空间精度)与音画相关性得分(评估时间对齐)。定量分析与用户研究均表明,本系统生成的动作质量极高,常与真实表演无法区分。

原文摘要 · Abstract (English)

Music-driven character animation enables and enhances transformative applications in entertainment and interactive education. However, synthesizing realistic drumming motion from audio remains challenging due to the inherent tension between high-acceleration dynamics and the need for extreme spatial-temporal precision. Existing approaches, often reliant on motion matching or MIDI input, struggle with generalizing to diverse real-world audio. Moreover, the field lacks standardized evaluation metrics capable of distinguishing precise drumming from noisy motion. In this paper, we introduce a generative diffusion framework featuring a dual-objective loss function that decouples skeletal integrity from drumstick precision, thus enabling centimeter-level stick precision without sacrificing natural body dynamics. Additionally, leveraging our own dataset and data augmentation strategy, the model generalizes to non-curated, in-the-wild audio. To rigorously evaluate performance, we propose two novel metrics: an impact-to-target distance to quantify spatial precision and an audio-motion correlation score to assess temporal alignment. Our quantitative analysis and user studies demonstrate that our system generates high-quality motion that is often indistinguishable from ground-truth performances.

音频驱动动作生成扩散模型舞蹈动画

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。