arXiv:2605.28160cs.AI2026-05中稿 · ICML

让语言模型按需调用视觉模块,提升多模态推理准确性

Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal Reasoning

论文配图:Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal Reasoning
图 1 · 摘自论文原文
  • 语言模型自主决定何时调用视觉模块获取证据
  • 零样本下多个基准测试均超越现有方法
  • 适合需要精准视觉证据的复杂推理任务

现有多模态推理方法主要采用两种范式:先将视觉输入转为文本再推理,或在统一的视觉-语言表示空间中端到端推理。前者依赖静态视觉转文本,易丢失细节;后者受联合优化和注意力机制影响,出现语言主导现象,削弱对视觉证据的忠实度。本文提出一种认知调度框架CSMR,让语言模型自主决策何时调用独立的视觉感知模块以获取任务相关的视觉证据。在多个多模态推理基准上的实验表明,CSMR在零样本设置下持续优于代表性基线方法。进一步分析证实,性能优势主要源于所提出的认知调度机制。

原文摘要 · Abstract (English)

Existing multimodal reasoning approaches predominantly follow two paradigms: converting visual inputs into text prior to reasoning, or performing end-to-end reasoning within a unified vision-language representation space. Despite their empirical progress, both paradigms suffer from fundamental structural limitations. The former relies on static visual-to-text conversion, which tends to compress and lose fine-grained visual details. The latter is prone to linguistic dominance induced by joint optimization and attention mechanisms, leading to systematically weakened faithfulness to visual evidence during reasoning. In this work, we argue that a central challenge is how and when visual evidence is introduced into the reasoning process. Motivated by this insight, we propose CSMR, a multimodal reasoning framework in which a language model controls the reasoning process by deciding when to invoke an independent visual perception module to acquire task-relevant visual evidence. Experiments across multiple multimodal reasoning benchmarks show that CSMR consistently outperforms representative baseline methods in accuracy under a zero-shot setting. Further experimental analysis confirms that these advantages primarily arise from the proposed cognitive scheduling mechanism.

多模态推理认知调度视觉证据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。