arXiv:2606.07533cs.CLcs.AI2026-06

为多模态大模型提供可解释性分析工具,揭示语音与文本如何共同影响模型决策。

Bridging Traditional Explainability Methods and Multimodal Multilingual Models: An XAI-Based Analysis

  • 将沙普利值扩展到多模态数据,把文本词元和音频片段视为协作特征
  • 提出频谱引导的语音对齐方法,解决音视频粒度不匹配问题
  • 开源工具包支持交互式可视化,适合多语言多模态研究者使用

多模态大语言模型(MLLMs)能有效融合文本与语音以理解复杂对话语境,但异构模态如何影响模型行为的内部机制仍不透明。尽管沙普利值(SV)在文本NLP中提供稳健的局部可解释性框架,其向多模态数据的扩展受限于跨通道依赖、复杂对话结构及高密度音频表示带来的计算瓶颈。本文形式化了多模态沙普利值框架,将离散文本词元与对齐音频段视为合作特征。为保障计算可行性,采用精确计算(低维输入)与采样近似策略——包括蒙特卡洛置换与奈曼最优分配的分层采样——在有限算力下最小化方差。针对模态粒度差异,提出频谱引导的语音对齐(SGPA)预处理方法,将高频音频流映射至可解释的词对齐段。贡献有二:一是开源模型无关的Python包与配套图形界面,用于多模态归因计算与交互可视化;二是基于VoiceBench与Infinity Instruct数据集的子集,在多种多语言场景下评估框架。实验表明,输入模态是归因波动的主要驱动因素,且标准句法重要性代理在多模态跨语言情境中常无法预测模型注意力。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) effectively integrate text and audio to interpret context in complex interactive dialogues. However, the internal mechanisms by which heterogeneous modalities influence model behavior remain opaque. While Shapley Values (SV) provide a robust, model-agnostic framework for local explainability in text-based NLP, their extension to multimodal data is hindered by cross-channel dependencies, intricate dialogue structures, and the prohibitive computational complexity of dense audio representations. In this work, we formalize a multimodal extension of the Shapley Value framework, treating discrete text tokens and aligned audio segments as cooperative features. To ensure computational feasibility, we deploy a suite of efficient estimation strategies: exact SV computation for low-dimensional inputs and sampling-based approximations - including Monte Carlo permutations and stratified sampling with Neyman-optimal allocation - to minimize variance under constrained computational budgets. To resolve the granularity mismatch between modalities, we propose Spectrogram-Guided Phonetic Alignment (SGPA), a novel preprocessing method that maps high-frequency audio streams to interpretable, word-aligned segments. Our contribution is twofold: first, we provide an open-source, model-agnostic Python package and a companion GUI for the computation and interactive visualization of multimodal attributions. Second, we evaluate our framework using curated subsets of the VoiceBench and Infinity Instruct datasets across diverse multilingual scenarios. Our experimental results reveal that input modality is a primary driver of attribution volatility and demonstrate that standard syntactic importance proxies often fail to predict model attention in multimodal, cross-lingual contexts.

可解释性多模态沙普利值语音对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。