arXiv:2605.21433cs.SD2026-05被引 1

控制数据与预训练条件,发现辅助分支对乐器音乐生成有隐性关键作用。

Instrumental Text-to-Music Generation with Auxiliary Conditioning Branches

  • 用扩散Transformer模型,仅以语言和音色作为辅助分支输入
  • 移除辅助分支后音频质量全面下降,参数重用效果微弱
  • 适合关注模型设计机制与训练稳定性的研究者

文本到音乐生成已取得显著进展,现代自回归与基于扩散的模型能从自然语言提示生成逼真音乐。然而,这些进展大多依赖大规模训练数据与外部预训练,难以在数据和预训练受限条件下评估设计选择的有效性。本文使用基于扩散Transformer的骨干网络,在仅含乐器的文本到音乐任务中,让辅助的歌词与音色分支仅接收退化的条件信号,通过受控消融实验发现:移除这些分支后,模型在AudioBox美学评分、大模型评判及人工主观评分(MOS)上均表现更差;将节省的参数重新投入模型深度也仅带来轻微提升。这表明辅助分支可能在训练过程中充当架构锚点,其作用超出显式条件内容本身。我们通过与外部基线对比,并提交至ICME 2026学术文本到音乐(ATTM)挑战赛,性能组排名第一,综合主观评分最高;效率组为决赛入围者,客观指标并列第二。

原文摘要 · Abstract (English)

Text-to-music generation has advanced rapidly, with modern autoregressive and diffusion-based models producing convincing music from natural-language prompts. However, much of this progress relies on large-scale training data and external pretraining, making it difficult to isolate which design choices remain effective when data and pretraining are controlled. We study this setting using a Diffusion Transformer backbone with lyric and timbre conditioning, adapted to an instrumental-only text-to-music task in which the auxiliary lyric and timbre branches receive only degenerate conditioning signals. Through controlled ablations, we find that models retrained without these branches score lower across AudioBox aesthetics, LLM-as-judge, and human MOS, and that reinvesting the saved parameters as additional DiT depth recovers only marginally. This suggests the auxiliary branches may act as training-time architectural anchors whose contribution goes beyond their explicit conditioning content. We validate the same model through comparisons with external instrumental baselines and through our submission to the ICME 2026 Academic Text-to-Music (ATTM) Grand Challenge, where our Performance submission ranked first under both the objective metrics and the subsequent organizer-administered MOS over 35 raters, attaining the highest overall MOS across all challenge submissions, while our Efficiency submission was a finalist that tied for second under the objective metrics.

文本到音乐扩散模型生成机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。