arXiv:2510.21111cs.CV2025-10NeurIPS被引 3

让AI像人一样主动探索环境来推理,突破静态视觉局限。

PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments

  • 设计交互式任务让AI通过动作主动收集信息
  • 在新基准上达到领先性能,信息获取效率提升显著
  • 适合研究具身智能与主动推理的学者

多模态大语言模型的视觉推理多限于静态、全可见场景,难以应对现实环境中因遮挡或视角受限导致的信息不全问题。人类则通过移动、观察和操作物体,以感知-推理-行动闭环方式主动获取信息。受此启发,我们提出主动视觉推理(AVR)任务,将推理扩展至部分可观测、可交互环境。AVR要求智能体:(1) 通过序列物理动作主动获取信息,(2) 融合多步观测进行连贯推理,(3) 根据动态视觉反馈调整决策。为严谨评估,我们构建了CLEVR-AVR仿真基准,包含多轮交互环境,用于测试推理准确性和信息获取效率。同时推出AVR-152k数据集,提供丰富的思维链(CoT)标注,涵盖不确定性识别、动作条件下的信息增益预测及信息最大化动作选择,支持高阶马尔可夫决策过程训练。基于此,我们开发了PhysVLM-AVR,在CLEVR-AVR、具身推理(OpenEQA、RoboVQA)和被动视觉推理(GeoMath、Geometry30K)任务中均达当前最优。分析显示,现有具身MLLM虽能察觉信息缺失,却难以通过交互主动获取并整合新信息,暴露出主动推理能力的根本缺陷。

原文摘要 · Abstract (English)

Visual reasoning in multimodal large language models (MLLMs) has primarily been studied in static, fully observable settings, limiting their effectiveness in real-world environments where information is often incomplete due to occlusion or limited field of view. Humans, in contrast, actively explore and interact with their environment-moving, examining, and manipulating objects-to gather information through a closed-loop process integrating perception, reasoning, and action. Inspired by this human capability, we introduce the Active Visual Reasoning (AVR) task, extending visual reasoning to partially observable, interactive environments. AVR necessitates agents to: (1) actively acquire information via sequential physical actions, (2) integrate observations across multiple steps for coherent reasoning, and (3) dynamically adjust decisions based on evolving visual feedback. To rigorously evaluate AVR, we introduce CLEVR-AVR, a simulation benchmark featuring multi-round interactive environments designed to assess both reasoning correctness and information-gathering efficiency. We present AVR-152k, a large-scale dataset that offers rich Chain-of-Thought (CoT) annotations detailing iterative reasoning for uncertainty identification, action-conditioned information gain prediction, and information-maximizing action selection, crucial for training agents in a higher-order Markov Decision Process. Building on this, we develop PhysVLM-AVR, an MLLM achieving state-of-the-art performance on CLEVR-AVR, embodied reasoning (OpenEQA, RoboVQA), and passive visual reasoning (GeoMath, Geometry30K). Our analysis also reveals that current embodied MLLMs, despite detecting information incompleteness, struggle to actively acquire and integrate new information through interaction, highlighting a fundamental gap in active reasoning capabilities.

具身智能主动推理视觉推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。