arXiv:2507.09904cs.SDeess.AS2025-07被引 2

用双分支模型自动评估音乐生成质量,准确率大幅超越基线。

ASTAR-NTU solution to AudioMOS Challenge 2025 Track1

  • 双分支结构融合预训练音频与文本模型,通过交叉注意力对齐
  • 将音乐评分转为软分布,提升对等级数据的建模能力
  • 单模型在测试集上达0.991(MI)和0.952(TA)相关性,显著领先

文本到音乐系统的评估受限于专家标注的成本与获取难度。AudioMOS 2025 Challenge Track 1旨在自动预测音乐印象(MI)及提示与生成音乐之间的文本对齐(TA)。本文报告了我们的获胜系统:采用基于预训练MuQ和RoBERTa模型的双分支架构作为音视频与文本编码器,通过交叉注意力融合表示。训练时将MI与TA预测重构为分类任务,并利用高斯核将独热标签转换为软分布以体现MOS分数的有序特性。在官方测试集上,单模型实现系统级Spearman秩相关系数(SRCC)0.991(MI)与0.952(TA),较挑战基线分别提升21.21%与31.47%。

原文摘要 · Abstract (English)

Evaluation of text-to-music systems is constrained by the cost and availability of collecting experts for assessment. AudioMOS 2025 Challenge track 1 is created to automatically predict music impression (MI) as well as text alignment (TA) between the prompt and the generated musical piece. This paper reports our winning system, which uses a dual-branch architecture with pre-trained MuQ and RoBERTa models as audio and text encoders. A cross-attention mechanism fuses the audio and text representations. For training, we reframe the MI and TA prediction as a classification task. To incorporate the ordinal nature of MOS scores, one-hot labels are converted to a soft distribution using a Gaussian kernel. On the official test set, a single model trained with this method achieves a system-level Spearman's Rank Correlation Coefficient (SRCC) of 0.991 for MI and 0.952 for TA, corresponding to a relative improvement of 21.21\% in MI SRCC and 31.47\% in TA SRCC over the challenge baseline.

音乐生成自动评估跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。