用可解释性方法分析音乐语音大模型的听觉依赖程度
Investigating Modality Contribution in Audio LLMs for Music
- 用Shapley值量化音频与文本对模型输出的贡献
- 高准确率模型更依赖文本,但音频仍能定位关键声音事件
- 首次将MM-SHAP用于音频大模型,推动可解释音频AI发展
音频大语言模型(Audio LLMs)使人类能够就音乐进行类人对话,但其是否真正聆听音频仍不明确。本文通过量化各模态对模型输出的贡献来探究此问题。我们采用基于Shapley值的无性能依赖评分框架MM-SHAP,评估两个模型在MuChoMusic基准上的表现。结果表明,准确率更高的模型更依赖文本回答问题;但进一步分析显示,即使整体音频贡献较低,模型仍能成功定位关键声学事件,说明音频信息并未被完全忽略。本研究首次将MM-SHAP应用于音频大模型,为可解释人工智能和音频理解研究奠定基础。
原文摘要 · Abstract (English)
Audio Large Language Models (Audio LLMs) enable human-like conversation about music, yet it is unclear if they are truly listening to the audio or just using textual reasoning, as recent benchmarks suggest. This paper investigates this issue by quantifying the contribution of each modality to a model's output. We adapt the MM-SHAP framework, a performance-agnostic score based on Shapley values that quantifies the relative contribution of each modality to a model's prediction. We evaluate two models on the MuChoMusic benchmark and find that the model with higher accuracy relies more on text to answer questions, but further inspection shows that even if the overall audio contribution is low, models can successfully localize key sound events, suggesting that audio is not entirely ignored. Our study is the first application of MM-SHAP to Audio LLMs and we hope it will serve as a foundational step for future research in explainable AI and audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。