arXiv:2509.17740cs.CVcs.CL2025-09EMNLP被引 7

用弱监督生成可解释推理链,提升多模态模型图像分类能力

WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification

  • 基于概念瓶颈模型重构推理链,弱监督生成解释
  • 跨10个数据集提升可解释性37%,并改善分类准确率
  • 适合需要细粒度视觉理解的多模态模型优化场景

多模态大语言模型(MLLMs)在视觉-文本推理中表现优异,多模态思维链(MCoT)提示显著提升了可解释性。然而现有MCoT方法依赖丰富的理由数据集,主要关注对象间推理,忽视了图像分类中至关重要的对象内理解。为此,我们提出WISE:一种弱监督引导的逐步解释方法,通过将概念瓶颈模型(CBMs)的概念表征重构为简洁、可解释的推理链,对任意图像分类数据集进行MCoT增强。在十个数据集上的实验表明,生成的MCoT不仅使可解释性提升37%,还将用于微调MLLMs时带来分类准确率提升。本工作连接了基于概念的可解释性与生成式MCoT推理,为提升多模态模型的细粒度视觉理解提供通用框架。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have shown promise in visual-textual reasoning, with Multimodal Chain-of-Thought (MCoT) prompting significantly enhancing interpretability. However, existing MCoT methods rely on rationale-rich datasets and largely focus on inter-object reasoning, overlooking the intra-object understanding crucial for image classification. To address this gap, we propose WISE, a Weak-supervision-guided Step-by-step Explanation method that augments any image classification dataset with MCoTs by reformulating the concept-based representations from Concept Bottleneck Models (CBMs) into concise, interpretable reasoning chains under weak supervision. Experiments across ten datasets show that our generated MCoTs not only improve interpretability by 37% but also lead to gains in classification accuracy when used to fine-tune MLLMs. Our work bridges concept-based interpretability and generative MCoT reasoning, providing a generalizable framework for enhancing MLLMs in fine-grained visual understanding.

多模态可解释性弱监督图像分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。