arXiv:2509.15625cs.SDeess.AS2025-09被引 5

用敲击声生成高保真鼓乐,无需训练即可适配多种音色。

The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling

  • 用掩码语言模型将节奏声音手势转为鼓乐音频。
  • 仅用不到10小时数据训练,零样本生成高质量鼓声。
  • 适合音乐创作、实时互动和跨音色迁移场景。

无论是音乐人还是非音乐人,常通过拍手、beatbox等节奏性声音手势表达鼓点创意。然而将这些创意转化为完整制作的鼓乐录音耗时较长,影响创作流程。为此,我们提出TRIA(The Rhythm In Anything),一种基于掩码变换器的模型,可将目标节奏的声音提示与代表鼓组音色的第二个提示,映射为符合节奏且带有合理润色的高保真鼓乐音频。主观与客观评估表明,仅使用不足10小时公开鼓乐数据训练的TRIA模型,可在零样本条件下生成跨多种音色的高质量节奏实现。

原文摘要 · Abstract (English)

Musicians and nonmusicians alike use rhythmic sound gestures, such as tapping and beatboxing, to express drum patterns. While these gestures effectively communicate musical ideas, realizing these ideas as fully-produced drum recordings can be time-consuming, potentially disrupting many creative workflows. To bridge this gap, we present TRIA (The Rhythm In Anything), a masked transformer model for mapping rhythmic sound gestures to high-fidelity drum recordings. Given an audio prompt of the desired rhythmic pattern and a second prompt to represent drumkit timbre, TRIA produces audio of a drumkit playing the desired rhythm (with appropriate elaborations) in the desired timbre. Subjective and objective evaluations show that a TRIA model trained on less than 10 hours of publicly-available drum data can generate high-quality, faithful realizations of sound gestures across a wide range of timbres in a zero-shot manner.

音频生成节奏建模零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。