arXiv:2602.01541cs.CVcs.AI2026-02被引 7

让多模态大模型像人一样‘脑内成像’,提升复杂推理能力。

Toward Cognitive Supersensing in Multimodal Large Language Model

  • 引入视觉意象预测头,让模型在脑中生成视觉化推理链。
  • 在认知评测基准上显著超越现有模型,跨领域泛化能力强。
  • 适合研究视觉推理、认知建模的学者和开发者参考。

多模态大语言模型(MLLM)在开放词汇感知任务中表现卓越,但在处理需抽象视觉细节与视觉记忆的复杂认知问题时仍受限。当前方法主要在文本空间扩展思维链(CoT)推理,忽视了类比人类视空间草图板的视觉推理机制。为此,我们提出认知超感(Cognitive Supersensing)训练范式,通过引入潜在视觉意象预测(LVIP)头,联合学习视觉认知潜在表示序列,并与答案对齐,形成基于视觉的内部推理链。进一步引入强化学习阶段,基于此具身化的视觉潜在优化文本推理路径。为评估模型的认知能力,我们构建了CogSense-Bench,一个涵盖五个认知维度的视觉问答(VQA)基准。大量实验表明,采用认知超感训练的MLLM在CogSense-Bench上显著优于先进基线,并在跨域数学与科学类VQA任务中展现出更强泛化能力,提示内在视觉意象可能是弥合感知识别与认知理解差距的关键。代码与模型权重将开源。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have achieved remarkable success in open-vocabulary perceptual tasks, yet their ability to solve complex cognitive problems remains limited, especially when visual details are abstract and require visual memory. Current approaches primarily scale Chain-of-Thought (CoT) reasoning in the text space, even when language alone is insufficient for clear and structured reasoning, and largely neglect visual reasoning mechanisms analogous to the human visuospatial sketchpad and visual imagery. To mitigate this deficiency, we introduce Cognitive Supersensing, a novel training paradigm that endows MLLMs with human-like visual imagery capabilities by integrating a Latent Visual Imagery Prediction (LVIP) head that jointly learns sequences of visual cognitive latent embeddings and aligns them with the answer, thereby forming vision-based internal reasoning chains. We further introduce a reinforcement learning stage that optimizes text reasoning paths based on this grounded visual latent. To evaluate the cognitive capabilities of MLLMs, we present CogSense-Bench, a comprehensive visual question answering (VQA) benchmark assessing five cognitive dimensions. Extensive experiments demonstrate that MLLMs trained with Cognitive Supersensing significantly outperform state-of-the-art baselines on CogSense-Bench and exhibit superior generalization on out-of-domain mathematics and science VQA benchmarks, suggesting that internal visual imagery is potentially key to bridging the gap between perceptual recognition and cognitive understanding. We will open-source the CogSense-Bench and our model weights.

多模态视觉推理认知建模大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。