融合视觉与文本解释,提升AI模型可解释性与分类准确率
MEGL: Multimodal Explanation-Guided Learning
- 用视觉热图引导文本解释,实现空间与语义双重定位
- 在无标注情况下仍能对齐图文解释,提升一致性
- 新数据集验证,兼顾解释质量与预测性能
解释人工智能模型的决策过程对缓解其'黑箱'问题至关重要,尤其在图像分类任务中。传统可解释AI方法多依赖单一模态解释:视觉解释虽能定位关键区域,但缺乏理由;文本解释提供上下文却无空间锚定。两者常不一致或不完整,影响可靠性。为此,我们提出多模态解释引导学习(MEGL)框架,融合视觉与文本解释以增强可解释性并提升分类性能。提出的显著性驱动文本定位(SDTG)方法将视觉显著图的空间信息融入文本解释,生成兼具空间定位与语义丰富性的解释。此外,引入文本监督机制对齐视觉解释与文本说明,即使在缺少真实视觉标注时亦可实现。进一步设计视觉解释分布一致性损失,使生成解释符合数据集整体模式,从而有效利用不完整的多模态监督信号。我们在两个新构建的数据集Object-ME和Action-ME上验证了MEGL,结果表明其在预测准确率与解释质量方面均优于现有方法,涵盖视觉与文本双维度表现。
原文摘要 · Abstract (English)
Explaining the decision-making processes of Artificial Intelligence (AI) models is crucial for addressing their "black box" nature, particularly in tasks like image classification. Traditional eXplainable AI (XAI) methods typically rely on unimodal explanations, either visual or textual, each with inherent limitations. Visual explanations highlight key regions but often lack rationale, while textual explanations provide context without spatial grounding. Further, both explanation types can be inconsistent or incomplete, limiting their reliability. To address these challenges, we propose a novel Multimodal Explanation-Guided Learning (MEGL) framework that leverages both visual and textual explanations to enhance model interpretability and improve classification performance. Our Saliency-Driven Textual Grounding (SDTG) approach integrates spatial information from visual explanations into textual rationales, providing spatially grounded and contextually rich explanations. Additionally, we introduce Textual Supervision on Visual Explanations to align visual explanations with textual rationales, even in cases where ground truth visual annotations are missing. A Visual Explanation Distribution Consistency loss further reinforces visual coherence by aligning the generated visual explanations with dataset-level patterns, enabling the model to effectively learn from incomplete multimodal supervision. We validate MEGL on two new datasets, Object-ME and Action-ME, for image classification with multimodal explanations. Experimental results demonstrate that MEGL outperforms previous approaches in prediction accuracy and explanation quality across both visual and textual domains. Our code will be made available upon the acceptance of the paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。