arXiv:2606.10010eess.AScs.AI2026-06中稿 · IEEE Signal Proces…

提出新方法提升文本生成音乐的自动评估效果

DeRA-MOS: Optimizing Text-to-Music Evaluation via Decoupled Listwise Ranking and Modality Alignment

  • 分步优化音乐印象与文本对齐,用列表排序损失增强相关性
  • 在MusicEval上使两项评价指标显著提升,排名相关性更强
  • 适合需要高效、准确评估生成音乐质量的研究者

文本到音乐(TTM)系统评估成本高昂,因音乐印象(MI)和文本对齐(TA)分数依赖人工均值评分(MOS)。现有自动MOS估计算法多采用点对点回归或分布分类训练,未直接优化基于排名的指标,且对跨模态一致性约束较弱。为此,本文提出DeRA-MOS,一种解耦的TTM评估优化框架:针对MI,引入批处理感知的列表排序损失,建模每个小批次内的相对顺序,更契合斯皮尔曼等级相关系数(SRCC)评估标准;针对TA,提出基于得分锚定的模态对齐损失,将人类评分映射至目标音文相似度,并在融合前正则化潜在空间。实验证明,该解耦框架有效缓解了点对点训练偏差与模态漂移问题,在MusicEval数据集上显著提升MI与TA的排名性能,建立了一套稳健的大规模TTM评估范式。

原文摘要 · Abstract (English)

Evaluating text-to-music (TTM) systems remains expensive because music impression (MI) and text alignment (TA) scores rely on human mean opinion scores (MOS). Most automatic MOS estimators are trained with point-wise regression or distributional classification. These objectives do not directly optimize rank-based metrics and provide weak geometric constraints for cross-modal coherence. To address these gaps, we propose DeRA-MOS, a decoupled optimization framework for TTM evaluation. For MI, we introduce a batch-aware listwise ranking loss that models relative order within each mini-batch and better aligns with evaluation based on Spearman's rank correlation coefficient (SRCC). For TA, we introduce a score-anchored modality alignment loss that maps human scores to target audio-text similarity and regularizes the latent space before fusion. By effectively mitigating the point-wise training mismatch and modality drift, experiments on MusicEval demonstrate that our decoupled framework yields substantial improvements in both MI and TA ranking metrics, establishing a robust paradigm for large-scale TTM evaluation.

文本生成音乐自动评估跨模态对齐排名学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。