arXiv:2603.11689cs.AI2026-03中稿 · ECCV

为大模型零样本任务提供可解释的逻辑验证通道

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks

  • 新增显式逻辑通道,通过概率推理验证隐式模型行为
  • 跨通道一致性率实现无标注下的模型选择与性能提升
  • 适用于需要可信推理的多模态应用,如医疗、安全领域

前沿多模态大语言模型在视觉-语言理解任务中表现优异,但常以黑箱方式部署于新任务。本文提出显式逻辑通道(Explicit Logic Channel),与黑箱模型并行运行,用于逻辑推理、验证与增强。该通道模拟人类推理,结合大模型、视觉特征提取器与概率推理,对显式视觉证据进行事实性、反事实性和关系性推理。引入一致性率(CR)实现跨通道验证与模型选择,无需真实标签。跨通道融合进一步提升零样本任务性能,增强可信度。在三个挑战性基准上,针对两类典型视觉-语言理解任务(MC-VQA 和 HC-REC),评估了来自四大主流家族的11个开源前沿多模态模型。系统实验表明,该方法有效提升模型可解释性与可信度。

原文摘要 · Abstract (English)

Frontier Multimodal Large Language Models (MLLMs) exhibit remarkable capabilities in Visual-Language Comprehension (VLC) tasks. However, they are often deployed as zero-shot solution to new tasks in a black-box manner. Validating and understanding the behavior of these models become important for application to new task. We propose an Explicit Logic Channel, in parallel with the black-box model channel, to perform explicit logical reasoning for model validation, selection and enhancement. The frontier MLLM, encapsulating latent vision-language knowledge, can be considered as an Implicit Logic Channel. The proposed Explicit Logic Channel, mimicking human logical reasoning, incorporates a LLM, a VFM, and logical reasoning with probabilistic inference for factual, counterfactual, and relational reasoning over the explicit visual evidence. A Consistency Rate (CR) is proposed for cross-channel validation and model selection, even without ground-truth annotations. Additionally, cross-channel integration further improves performance in zero-shot tasks over MLLMs, grounded with explicit visual evidence to enhance trustworthiness. Comprehensive experiments conducted for two representative VLC tasks, i.e., MC-VQA and HC-REC, on three challenging benchmarks, with 11 recent open-source MLLMs from 4 frontier families. Our systematic evaluations demonstrate the effectiveness of proposed ELC and CR for model validation, selection and improvement on MLLMs with enhanced explainability and trustworthiness.

多模态模型逻辑推理零样本可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。