提出新方法实现音乐生成低延迟,速度提升30%且质量几乎不变。
SAGE-Music: Low-Latency Symbolic Music Generation via Attribute-Specialized Key-Value Head Sharing
- 设计属性专精的键值头共享机制,适配符号化音乐结构。
- 在多轨音乐生成中实现约30%推理加速,质量仅降0.4%。
- 适合需要实时交互的音乐创作场景,如人机协同即兴演奏。
低延迟符号化音乐生成对实时即兴演奏和人机共创至关重要。现有基于Transformer的模型在推理速度与音乐质量间存在权衡:传统加速技术如嵌入池化会显著降低质量,而近期提出的字节对编码(BPE)方法虽在单轨钢琴数据上有效,但在多轨设置下性能大幅下降,我们的分析揭示了这一问题。为此,我们提出属性专精的键值头共享(AS-KVHS),针对音乐的结构化符号表示进行优化,在客观评估中实现约30%的推理速度提升,质量仅轻微下降约0.4%,主观听感测试甚至略有改善。主要贡献包括:(1) 首次系统研究BPE在多轨符号化音乐中的泛化能力;(2) 提出AS-KVHS以实现低延迟符号化音乐生成。此外,我们还发布了SAGE-Music——一个开源基准,其生成质量达到或超越当前最先进水平。
原文摘要 · Abstract (English)
Low-latency symbolic music generation is essential for real-time improvisation and human-AI co-creation. Existing transformer-based models, however, face a trade-off between inference speed and musical quality. Traditional acceleration techniques such as embedding pooling significantly degrade quality, while recently proposed Byte Pair Encoding (BPE) methods - though effective on single-track piano data - suffer large performance drops in multi-track settings, as revealed by our analysis. We propose Attribute-Specialized Key-Value Head Sharing (AS-KVHS), adapted to music's structured symbolic representation, achieving about 30% inference speedup with only a negligible (about 0.4%) quality drop in objective evaluations and slight improvements in subjective listening tests. Our main contributions are (1) the first systematic study of BPE's generalizability in multi-track symbolic music, and (2) the introduction of AS-KVHS for low-latency symbolic music generation. Beyond these, we also release SAGE-Music, an open-source benchmark that matches or surpasses state-of-the-art models in generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。