构建细粒度音乐质量评估基准,助力生成更专业的歌曲。
SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment
- 从7个维度细评歌曲质量,覆盖演唱、旋律、编排等专业层面。
- 基于1.17万条专业人士标注数据,验证评估结果与专家评分高度一致。
- 可定位当前生成模型短板,适合音乐生成研究者使用。
近年来,文本转歌曲生成技术已能生成逼真的音乐内容,但现有评估基准缺乏捕捉多维审美细节的专业精细度。本文提出SongBench,一个面向七个关键维度(演唱、乐器、旋律、结构、编排、混音、音乐性)的细粒度歌曲质量评估框架。基于该框架,我们构建了一个由音乐专业人士标注的数据库,包含来自前沿模型的11,717个样本。实验结果表明,SongBench与专家评分具有高度相关性。通过揭示当前先进模型在各维度上的细微差距,SongBench可作为诊断工具,引导生成系统向更专业、更连贯的音乐创作方向发展。
原文摘要 · Abstract (English)
Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture multi-dimensional aesthetic nuances. In this paper, we propose SongBench, a specialized framework for fine-grained song assessment across seven key dimensions: Vocal, Instrument, Melody, Structure, Arrangement, Mixing, and Musicality. Utilizing this framework, we construct an expert-annotated database comprising 11,717 samples from state-of-the-art models, labeled by music professionals. Extensive experimental results demonstrate that SongBench achieves high correlation with expert ratings. By revealing fine-grained performance gaps in current state-of-the-art models, SongBench serves as a diagnostic benchmark to steer the development toward more professional and musically coherent song generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。