用文字和乐谱控制,生成自然有表现力的多乐器演奏音频。
RenderBox: Expressive Performance Rendering with Text Control
- 基于扩散Transformer与跨注意力机制,融合文本与乐谱实现双控生成。
- 在FAD、CLAP、节奏与音高准确率上优于基线模型。
- 适合音乐生成、智能作曲及交互式表演系统研究者使用。
富有表现力的音乐演奏生成旨在通过节奏、力度、连奏和乐器特有技巧等变化,诠释符号化乐谱并传达音乐情感意图。我们提出RenderBox,一个统一的跨乐器文本与乐谱可控音频演奏生成框架,通过自然语言进行粗粒度控制,利用乐谱实现细粒度调控。基于扩散Transformer架构与交叉注意力联合条件建模,设计了从纯合成到表现性演奏的渐进式训练范式,逐步引入速度、失误和风格多样性等可控因素。在多个关键指标上,如FAD和CLAP得分,以及不同提示任务下的节拍与音高准确性,RenderBox均优于基线模型。主观评估进一步表明,其生成的演奏具有自然感与音乐感染力,能有效响应提示内容与创作意图。
原文摘要 · Abstract (English)
Expressive music performance rendering involves interpreting symbolic scores with variations in timing, dynamics, articulation, and instrument-specific techniques, resulting in performances that capture musical can emotional intent. We introduce RenderBox, a unified framework for text-and-score controlled audio performance generation across multiple instruments, applying coarse-level controls through natural language descriptions and granular-level controls using music scores. Based on a diffusion transformer architecture and cross-attention joint conditioning, we propose a curriculum-based paradigm that trains from plain synthesis to expressive performance, gradually incorporating controllable factors such as speed, mistakes, and style diversity. RenderBox achieves high performance compared to baseline models across key metrics such as FAD and CLAP, and also tempo and pitch accuracy under different prompting tasks. Subjective evaluation further demonstrates that RenderBox is able to generate controllable expressive performances that sound natural and musically engaging, aligning well with prompts and intent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。