测试大模型主动找信息的能力,发现现实场景下表现远差于理论测试。
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
- 让模型从候选图中主动选图补全信息,模拟真实世界推理。
- 20个主流模型在主动推理任务上表现远低于被动推理,差距显著。
- 小模型靠视觉改进提升明显,大模型则需优化思考策略。
多模态大语言模型(MLLMs)在多项基准测试中表现出色,但现有评估大多基于被动推理——即在信息完整的前提下逐步推理解题。这与真实场景脱节,因为‘看到’并不等于‘理解’。核心问题在于:当信息不完整时,MLLM能否主动获取缺失证据?为填补这一空白,我们要求模型在无特定任务先验的情况下,从候选图像池中选择目标图像,以迭代方式完善决策。为此,我们提出GuessBench基准,包含感知导向与知识导向两类图像,用于系统评估主动推理能力。我们评测了20个先进MLLMs,发现其在主动推理上的表现远落后于被动设置,表明仍有巨大提升空间。细粒度分析揭示,精细感知与及时决策是关键挑战。消融实验显示,感知增强对小型模型有效,而思维导向方法则在所有模型尺寸上均带来稳定提升。这些结果为未来多模态主动推理研究指明方向。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have shown strong capabilities across a broad range of benchmarks. However, most existing evaluations focus on passive inference, where models perform step-by-step reasoning under complete information. This setup is misaligned with real-world use, where seeing is not enough. This raises a fundamental question: Can MLLMs actively acquire missing evidence under incomplete information? To bridge this gap, we require the MLLMs to actively acquire missing evidence and iteratively refine decisions under incomplete information, by selecting a target image from a candidate pool without task-specific priors. To support systematic study, we propose GuessBench, a benchmark with both perception-oriented and knowledge-oriented images for evaluating active reasoning in MLLMs. We evaluate 20 superior MLLMs and find that performance on active reasoning lags far behind it on passive settings, indicating substantial room for improvement. Further analysis identifies fine-grained perception and timely decision-making as key challenges. Ablation studies show that perceptual enhancements benefit smaller models, whereas thinking-oriented methods provide consistent gains across model sizes. These results suggest promising directions for future research on multimodal active reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。