传统评估指标难以反映音乐标题的语义准确性。
Do Captioning Metrics Reflect Music Semantic Alignment?
- 发现BLEU、METEOR等指标对语法变化敏感
- 人工评价与这些指标相关性低
- 呼吁重新设计音乐标题评估方法
音乐标题生成已成为一个有前景的研究方向,得益于先进语言生成模型的发展。然而,当前音乐标题的评估仍严重依赖于为其他领域设计的传统指标,如BLEU、METEOR和ROUGE,缺乏在音乐语境下的合理性依据。本文通过实例表明,这些指标对语法变化敏感,且与人类判断的相关性较差。我们旨在强调需对音乐标题生成的评估方式开展批判性重审。
原文摘要 · Abstract (English)
Music captioning has emerged as a promising task, fueled by the advent of advanced language generation models. However, the evaluation of music captioning relies heavily on traditional metrics such as BLEU, METEOR, and ROUGE which were developed for other domains, without proper justification for their use in this new field. We present cases where traditional metrics are vulnerable to syntactic changes, and show they do not correlate well with human judgments. By addressing these issues, we aim to emphasize the need for a critical reevaluation of how music captions are assessed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。