LEMON高效生成多模态模型的局部解释,能区分不同模态贡献。
LEMON: Local Explanations via Modality-aware OptimizatioN
- 用结构化稀疏代理模型捕捉多模态输入特征
- 相比基线减少35-67倍黑盒评估次数,提速2-8倍
- 适合需要快速、可解释多模态决策的场景
多模态模型广泛应用,但现有可解释性方法常为单模态、依赖架构或计算成本过高。我们提出LEMON(基于模态感知优化的局部解释),一种无需模型结构信息的多模态预测解释框架。LEMON通过拟合一个具有组结构稀疏性的模态感知代理模型,生成统一解释,分离出模态级贡献与特征级归因。该方法将预测器视为黑箱,计算高效,仅需少量前向传播,在重复扰动下仍保持高保真度。我们在视觉语言问答和包含图像、文本、表格输入的临床预测任务上评估LEMON,对比代表性多模态基线。在不同骨干网络下,LEMON在删除法保真度上表现相当,但黑盒评估次数减少35-67倍,运行时间缩短2-8倍。
原文摘要 · Abstract (English)
Multimodal models are ubiquitous, yet existing explainability methods are often single-modal, architecture-dependent, or too computationally expensive to run at scale. We introduce LEMON (Local Explanations via Modality-aware OptimizatioN), a model-agnostic framework for local explanations of multimodal predictions. LEMON fits a single modality-aware surrogate with group-structured sparsity to produce unified explanations that disentangle modality-level contributions and feature-level attributions. The approach treats the predictor as a black box and is computationally efficient, requiring relatively few forward passes while remaining faithful under repeated perturbations. We evaluate LEMON on vision-language question answering and a clinical prediction task with image, text, and tabular inputs, comparing against representative multimodal baselines. Across backbones, LEMON achieves competitive deletion-based faithfulness while reducing black-box evaluations by 35-67 times and runtime by 2-8 times compared to strong multimodal baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。