用少量数据生成高质量伴奏,精准控制音乐结构与语义。
S2Accompanist: A Semantic-Aware and Structure-Guided Diffusion Model for Music Accompaniment Generation

- 基于结构引导的扩散模型,融合语义标注与乐谱结构信息。
- 仅402万参数即达领先性能,效率赛道排名第一。
- 适合资源受限场景下的音乐生成,尤其关注伴奏质量与可控性。
高保真文本到音乐生成通常依赖大规模专有数据集和巨大计算资源。现有模型在生成连贯纯伴奏方面表现不佳,且因依赖粗粒度轨道级注释而缺乏精确的局部语义控制。为在数据与算力受限条件下解决上述问题,我们提出S2Accompanist,一种针对ICME2026 ATTM Grand Challenge设计的语义感知、结构引导扩散模型。我们构建了自动化数据流水线:包括结构分割、基于大音频-语言模型的段落级标题生成、双指标质量评估,以弥补原始数据中缺乏局部元数据的问题。此外,我们提出一种语义感知的变分自编码器微调策略,将基础乐谱结构(LeadSheet)显式地融入声学潜在空间,显著提升音频整体保真度。大量实验表明,S2Accompanist在ATTM Grand Challenge基准上,于效率与性能双赛道均达到当前最优客观性能。仅含402M参数,其表现仍优于更大规模的无约束模型,并在效率赛道中夺得第一。
原文摘要 · Abstract (English)
High-fidelity text-to-music generation typically relies on massive proprietary datasets and immense computational resources. Existing models often struggle to generate coherent pure musical accompaniments and lack precise, localized semantic control due to their reliance on coarse, track-level annotations. To address these limitations under constrained data and computing resources, we propose S2Accompanist, a Semantic-Aware and Structure-Guided Diffusion Model developed for the ICME2026 ATTM Grand Challenge. Specifically, we design an automated data pipeline comprising structural segmentation, Large Audio-Language Model driven segment-level captioning, and dual-metric quality grading to overcome the absence of localized metadata in raw datasets. Furthermore, we propose a semantic-aware Variational Autoencoder fine-tuning strategy that explicitly distills foundational LeadSheet structures into the acoustic latent space, effectively improving the overall audio fidelity. Extensive experiments demonstrate that S2Accompanist achieves state-of-the-art objective performance on the ATTM Grand Challenge benchmark across both the Efficiency and Performance Tracks. With only 402M parameters, our model remains competitive compared to larger-scale unconstrained models and secured first place in the Efficiency Track.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。