首个评估歌声指导型语音-语言模型的基准,推动音频理解从描述迈向专业反馈。
VocalCoachBench: Benchmarking Audio-Language Models on Expert Feedback for Singing

- 构建包含12051条诊断主张的歌声教练评测集,区分结构化目标与自由表述反馈
- 12个模型在细粒度问题识别上低于标签先验水平,严格诊断对齐率不足7%
- 适合研究音频-语言模型在音乐教育、专业辅导等需精准建议场景的应用
近期语音-语言模型多聚焦于音频的识别、描述与推理,但专家级应用需要能基于输入生成指出问题并提出改进方案的反馈。我们提出VocalCoachBench,一个用于评估语音-语言模型在歌唱教学中生成专家级反馈的基准。该数据集包含515段录音,由18位专业声乐教练标注,共产生1056份专家反馈和12051条原子级教学主张。数据集包含同歌子集(用于控制对比)和多样化歌曲子集(支持跨曲目、多条件的片段级反馈)。为适应专家反馈的开放性,基准将确定性结构化目标与基于主张的评估分离。人类标注分析显示,专家一致性随标签粒度显著变化,由此催生分层结构化指标与主张式评估方法。对12个近期语音-语言模型的实验表明:模型虽能比较演唱表现并识别宽泛问题领域,但顶级3个细粒度问题标签识别率仍低于标签先验水平,严格诊断对齐率低于7%。据我们所知,VocalCoachBench是首个公开的歌唱专家反馈评测平台,推动语音-语言模型评价从描述走向分析性反馈。
原文摘要 · Abstract (English)
Recent audio-language models are increasingly evaluated on recognizing, describing, and reasoning about audio, but expert-facing applications require a different capability: producing feedback that identifies problems and suggests corrective actions grounded in the input. We introduce VocalCoachBench, a benchmark for evaluating audio-language models on expert vocal coaching feedback for singing. VocalCoachBench contains 515 recordings annotated by 18 professional vocal trainers, yielding 1,056 expert submissions and 12,051 atomic coaching claims. It comprises a same-song subset for controlled comparison and a diverse-song subset for segment-grounded feedback across varied songs and recording conditions. To accommodate the open-ended nature of expert feedback, VocalCoachBench sep- arates deterministic structured targets from claim-based assessment of free-form diagnosis and corrective guidance. Human annotation analysis shows that expert agreement varies strongly with label granularity, motivating hierarchical structured metrics and claim-based evaluation of open-ended feedback. Experiments with 12 recent audio-language models reveal a consistent gap: while models can compare performances and identify broad issue domains in free-form feedback, Top-3 fine-grained issue-label identification remains below label-prior baselines and strict diagnosis alignment stays below 7%. To our knowledge, VocalCoachBench pro- vides the first public testbed for evaluating audio-grounded expert feedback for singing, moving audio-language evaluation beyond description toward analytic feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。