arXiv:2607.10796cs.CV2026-07

让视觉语言模型像人一样分步思考,减少幻觉。

Mixture of Cognitive Experts in Large Vision-Language Models

论文配图:Mixture of Cognitive Experts in Large Vision-Language Models
图 1 · 摘自论文原文
  • 用认知分类框架分层调度多专家视觉模型输出
  • 推理过程可追踪,证据使用量提升37%
  • 适合需要可解释性的医疗、金融等高风险场景

大型视觉语言模型需对视觉与文本输入进行强推理。近期研究指出,认知要素如多样化表征与元认知与性能正相关。诸多感知功能已有特定领域计算机视觉模型提供,作为检测物体、定位、推断状态、恢复空间布局和读取文本的感知子系统。关键挑战在于将这些多编码器专家整合为可信、可解释且连贯的表示,以提升可验证性并减少幻觉。这很困难,因为视觉语言问题涵盖不同认知层级,而现有流水线对所有查询均采用相同感知-推理路径。我们提出一种基于布卢姆分类法的证据驱动多模态推理框架。两阶段认知语义化首先将专家输出分解为短小原子证据陈述,生成字面证据摘要;随后进行布卢姆语义化,将这些证据项转化为分阶段推理轨迹,轻量级推理轨迹模块定量分析该轨迹,使证据使用与推理进展显式可见。通过此集成,观察到感知与推理能力显著提升。此外,轨迹模块提供了量化证据:不同查询引发不同认知入口层级与证据使用路径,支持细粒度分析。

原文摘要 · Abstract (English)

Large Vision Language Models (LVLMs) require strong reasoning over both visual and textual input. Recent work suggests that cognitive elements, especially diverse representations and metacognition, correlate with better performance. Many of the needed perceptual functions are already provided by specialized domain-specific computer vision models, which act as the perceptual subsystem for detecting objects, localizing them, inferring states, recovering spatial layout, and reading text. The key challenge is to integrate these multi-encoder experts into a trustworthy, interpretable, and coherent representation that improves verifiability and reduces hallucinations. This is difficult because vision-language questions span different cognitive levels, yet most LVLM pipelines apply the same perception-reasoning routing regardless of the demand of each query. We propose an evidence-driven multimodal reasoning framework that utilizes a Bloom-inspired taxonomy as a hierarchical reasoning protocol. The two-stage cognitive verbalization first produces a Literal Evidence Summary by decomposing expert outputs into short, atomic evidence statements. It then performs Bloom Verbalization to turn these evidence items into a staged reasoning trace, and a lightweight Reasoning Trace Module quantitatively analyzes the trace to make evidence usage and reasoning progression explicit. Through this integration, we observed several improvements in perception and reasoning abilities. Moreover, the trace module provides quantitative evidence that different queries induce different cognitive entry levels and evidence-use trajectories that enable fine-grained analysis.

视觉语言模型认知架构可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。