梳理音乐生成评估体系,揭示现有方法缺陷并指明改进方向
A Survey on Evaluation Metrics for Music Generation
- 构建音频与符号音乐的评估指标分类体系
- 指出客观指标与人类感知相关性差、跨文化偏见等问题
- 适合音乐生成研究者与评估框架设计者参考
尽管音乐生成系统取得显著进展,但评估方法因音乐结构、连贯性、创造性及情感表达等复杂特性而未同步发展。本文揭示这一研究缺口,提出针对音频与符号音乐表示的评估指标详细分类体系,并进行批判性回顾,指出当前方法存在客观指标与人类感知相关性低、跨文化偏差及缺乏标准化等问题,阻碍模型间比较。为此,本文进一步提出未来研究方向,旨在构建全面的音乐生成评估框架。
原文摘要 · Abstract (English)
Despite significant advancements in music generation systems, the methodologies for evaluating generated music have not progressed as expected due to the complex nature of music, with aspects such as structure, coherence, creativity, and emotional expressiveness. In this paper, we shed light on this research gap, introducing a detailed taxonomy for evaluation metrics for both audio and symbolic music representations. We include a critical review identifying major limitations in current evaluation methodologies which includes poor correlation between objective metrics and human perception, cross-cultural bias, and lack of standardization that hinders cross-model comparisons. Addressing these gaps, we further propose future research directions towards building a comprehensive evaluation framework for music generation evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。