现有音乐问答评测骗过了模型,新框架让它们真听懂音乐。
Are you really listening? Boosting Perceptual Awareness in Music-QA Benchmarks
- 用文本模型的输出概率设计感知指数,量化问题对音频的依赖度。
- 在MuchoMusic上,纯文本模型得分从56.4%降至随机水平,验证评测有效性。
- 适合关注音频理解真实能力评估的研究者和开发者。
大型音频语言模型(LALMs)在音乐理解方面取得显著进展,但当前评估方法存在严重缺陷:在主流音乐问答基准MuchoMusic上,无音频感知能力的纯文本大模型最高可达56.4%准确率,与多数LALMs相当甚至更高。此外,当输入为随机高斯噪声时,LALMs仍显著高于随机水平。这表明现有基准主要衡量推理能力而非音频感知。为此,我们提出鲁棒听觉理解框架RUListening,引入感知指数(PI),通过分析文本模型对音频的对数概率分布,量化问题对音频感知的依赖程度。基于该指标生成合成难例,构建需真实音频感知的问答对。应用于MuchoMusic后,纯文本模型性能降至随机水平,而LALMs在音频替换为噪声时同样显著下降。结果验证了该框架有效提升评估对音频感知的真实要求。
原文摘要 · Abstract (English)
Large Audio Language Models (LALMs), where pretrained text LLMs are finetuned with audio input, have made remarkable progress in music understanding. However, current evaluation methodologies exhibit critical limitations: on the leading Music Question Answering benchmark, MuchoMusic, text-only LLMs without audio perception capabilities achieve surprisingly high accuracy of up to 56.4%, on par or above most LALMs. Furthermore, when presented with random Gaussian noise instead of actual audio, LALMs still perform significantly above chance. These findings suggest existing benchmarks predominantly assess reasoning abilities rather than audio perception. To overcome this challenge, we present RUListening: Robust Understanding through Listening, a framework that enhances perceptual evaluation in Music-QA benchmarks. We introduce the Perceptual Index (PI), a quantitative metric that measures a question's reliance on audio perception by analyzing log probability distributions from text-only language models. Using this metric, we generate synthetic, challenging distractors to create QA pairs that necessitate genuine audio perception. When applied to MuchoMusic, our filtered dataset successfully forces models to rely on perceptual information-text-only LLMs perform at chance levels, while LALMs similarly deteriorate when audio inputs are replaced with noise. These results validate our framework's effectiveness in creating benchmarks that more accurately evaluate audio perception capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。