对比自动评估与人类偏好,发现音乐生成模型评价存在明显偏差
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
- 用多维度感知评估和参考指标对比五种顶尖音乐生成模型
- 发现自动指标与人类偏好差异显著,现有评估方法存在局限
- 开源多模型音乐样本数据集,推动更以人为中心的评价研究
生成模型的评估仍是核心挑战,尤其当目标是反映人类偏好时。本文以音乐生成为案例,探究自动评估指标与人类偏好之间的差距。我们对五种前沿音乐生成方法进行了对比实验,从感知质量与人类创作音乐的分布相似性两方面进行评估。具体包括多个感知维度的合成音乐评价,以及基于参考的指标如Mauve Audio Divergence(MAD)和Kernel Audio Distance(KAD)。结果揭示了不同指标间存在显著不一致,凸显当前评估实践的局限性。为支持后续研究,我们发布了一个包含多个模型生成样本的基准数据集。本研究为生成模型中人类偏好对齐提供了更广阔的视角,倡导跨领域采用更以人为本的评估策略。
原文摘要 · Abstract (English)
Evaluating generative models remains a fundamental challenge, particularly when the goal is to reflect human preferences. In this paper, we use music generation as a case study to investigate the gap between automatic evaluation metrics and human preferences. We conduct comparative experiments across five state-of-the-art music generation approaches, assessing both perceptual quality and distributional similarity to human-composed music. Specifically, we evaluate synthesis music from various perceptual dimensions and examine reference-based metrics such as Mauve Audio Divergence (MAD) and Kernel Audio Distance (KAD). Our findings reveal significant inconsistencies across the different metrics, highlighting the limitation of the current evaluation practice. To support further research, we release a benchmark dataset comprising samples from multiple models. This study provides a broader perspective on the alignment of human preference in generative modeling, advocating for more human-centered evaluation strategies across domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。