arXiv:2605.28215cs.AIcs.CL2026-05中稿 · ICML

探究MLLM在少样本图像分类中解释能力,发现解释比预测更难。

Explaining is Harder Than Predicting Alone: Evaluating Concept-based Explanations of MLLMs as ICL Visual Classifiers

论文配图:Explaining is Harder Than Predicting Alone: Evaluating Concept-based Explanations of MLLMs as ICL Visual Classifiers
图 1 · 摘自论文原文
  • 通过五种严谨程度递增的条件评估模型概念解释能力。
  • 强制生成形式化解释导致准确率从93.8%降至90.1%。
  • 能正确描述判别性视觉特征的模型预测更准确,适合可解释性研究者。

上下文学习(ICL)使多模态大语言模型(MLLM)能够基于少量标注样例进行图像分类。然而,这些模型如何利用所提供上下文仍不透明。尽管思维链提示被广泛使用,但近期研究认为其可能无法反映真实的内部计算过程。本文系统评估了冻结的MLLM在少样本ICL下的概念可解释性,采用五种严谨程度递增的评估条件,从基础分类到描述逻辑(DL)公理生成。通过独立的LLM作为裁判管道评估四种前沿MLLM,我们证明:解释确实比预测更困难。令人惊讶的是,强制模型生成形式化、概念化的解释会单调降低预测准确率(从93.8%降至90.1%),与‘显式推理普遍提升性能’的假设相悖。然而,当模型成功阐述类别判别性视觉特征时,解释质量与正确预测强相关。结果表明,尽管MLLM擅长视觉分类,但缺乏实现形式化、机器可验证解释所需的特定指令微调。

原文摘要 · Abstract (English)

In-context learning (ICL) enables multimodal large language models (MLLMs) to classify images from a few labelled examples. Yet, how these models use the provided context remains opaque. While Chain-of-Thought prompting is widely used, recent work argues that it may not reflect true internal computation. In this paper, we systematically evaluate the concept-based explainability of frozen MLLMs under few-shot ICL using five conditions of increasing formal rigour, ranging from baseline classification to Description Logics (DL) axiom generation. Evaluating four state-of-the-art MLLMs via an independent LLM-as-a-judge pipeline, we demonstrate that explaining is genuinely harder than predicting alone. Surprisingly, forcing models to generate formally structured, concept-based explanations degrades predictive accuracy monotonically (from 93.8% to 90.1%), contradicting the assumption that explicit reasoning universally aids performance. However, when models successfully articulate class-discriminative visual features, explanation quality strongly correlates with correct predictions. Our findings suggest that while MLLMs excel at visual classification, they lack the specific instruction-tuning required for formal, machine-verifiable explainability.

多模态模型可解释性少样本学习推理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。