arXiv:2606.07387cs.LG2026-06被引 1

用音频对齐分数指导训练,让小数据也能生成高质量音乐。

Making the Most of Limited Data: Score-Aware Training for Text-to-Music Generation

论文配图:Making the Most of Limited Data: Score-Aware Training for Text-to-Music Generation
图 1 · 摘自论文原文
  • 用音频-文本匹配分做全程监督信号,低分片段改造成正则化训练数据。
  • 在ICME 2026挑战赛中,450M参数模型双项排名前3,效率与质量兼优。
  • 适合资源有限但想提升文本到音乐生成效果的研究者与开发者。

当前最先进的文本到音乐生成系统依赖海量专有数据和工业级算力,难以区分模型架构贡献与资源优势。我们提出「得分感知训练」,将音频-文本对齐得分作为全流程直接监督信号。不丢弃低分片段,而是通过CLAP条件化的贝塔噪声时间表将其导向高噪声训练阶段,发挥隐式正则化作用。同时,段级过滤剔除最不匹配样本,两阶段标题生成流程弥合了训练时的冗长描述与推理时简洁提示之间的分布差距。额外引入REPA辅助损失,无需额外数据即可从预训练的CLAP和MuQ编码器迁移结构化语义知识。基于FluxAudio的450M参数系统参与了ICME 2026 ATTM Grand Challenge效率赛道,客观评估中两项均获第2名,最终主观评分(MOS)在效率赛道位列第3。

原文摘要 · Abstract (English)

State-of-the-art text-to-music generation systems rely on massive proprietary datasets and industrial-scale compute, making it impossible to disentangle architectural contributions from resource advantages. We propose \textit{score-aware training}, which treats audio-caption alignment score as a direct supervision signal throughout the pipeline. Rather than discarding low-scoring segments, we repurpose them via a CLAP-conditioned Beta noise timestep schedule that routes them to high-noise training regimes, acting as an effective implicit regularizer. Complementarily, segment-level filtering removes the most misaligned examples, and a two-stage caption procedure bridges the distribution gap between verbose training captions and concise inference prompts. A REPA auxiliary loss further transfers structured semantic knowledge from pretrained CLAP and MuQ encoders without additional data. Our 450M-parameter FluxAudio-based system, submitted to the ICME 2026 ATTM Grand Challenge Efficiency Track, ranked 2nd across both tracks in the objective evaluation and 3rd in the Efficiency Track in the final MOS evaluation.

文本到音乐小数据训练得分感知高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。