让音乐多模态模型的决策过程可解释,看清音频与歌词如何协同影响结果。
MusicLIME: Explainable Multimodal Music Understanding
- 提出跨模态特征重要性分析方法,捕捉音频与歌词的交互作用。
- 通过局部解释聚合生成全局行为图谱,揭示模型整体决策模式。
- 适合关注音乐理解模型公平性与透明度的研究者与开发者。
多模态模型在音乐理解任务中至关重要,因它们能捕捉音频与歌词之间的复杂互动。然而,随着这些模型日益普及,可解释性需求随之增长——理解系统决策机制对于确保公平性、减少偏见和建立信任至关重要。本文提出 MusicLIME,一种适用于多模态音乐模型的模型无关特征重要性解释方法。不同于传统单模态方法(仅分别分析音频或歌词,忽略二者交互,常导致不完整或误导性解释),MusicLIME 揭示音频与歌词特征如何共同贡献于预测,提供模型决策的全景视图。此外,我们通过聚合局部解释生成全局解释,帮助用户获得模型行为的宏观视角。本工作推动了多模态音乐模型的可解释性提升,赋能用户做出明智选择,促进更公平、公正与透明的音乐理解系统发展。
原文摘要 · Abstract (English)
Multimodal models are critical for music understanding tasks, as they capture the complex interplay between audio and lyrics. However, as these models become more prevalent, the need for explainability grows-understanding how these systems make decisions is vital for ensuring fairness, reducing bias, and fostering trust. In this paper, we introduce MusicLIME, a model-agnostic feature importance explanation method designed for multimodal music models. Unlike traditional unimodal methods, which analyze each modality separately without considering the interaction between them, often leading to incomplete or misleading explanations, MusicLIME reveals how audio and lyrical features interact and contribute to predictions, providing a holistic view of the model's decision-making. Additionally, we enhance local explanations by aggregating them into global explanations, giving users a broader perspective of model behavior. Through this work, we contribute to improving the interpretability of multimodal music models, empowering users to make informed choices, and fostering more equitable, fair, and transparent music understanding systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。