让AI像人一样看图、操作、推理,解决空间认知难题
Perceive, Interact, Reason: Building Tool-Augmented Visual Agents for Spatial Reasoning

- 用视觉感知和交互工具主动获取空间信息
- 在13个任务上比基线模型提升4.4%至14.8%
- 适合需要精细空间理解的智能体应用
尽管近期视觉语言模型(VLMs)展现出强大的多模态理解能力,但在需要主动获取证据和多步视觉交互的空间推理任务中仍显不足。这表明仅依赖视觉编码器的隐式表征不足以恢复细粒度空间证据。本文提出PERIA——一种用于地图推理、视觉探查和视觉重建任务的工具增强型视觉智能体。PERIA采用两类轻量级工具:视觉感知工具用于揭示文本、符号和空间证据;视觉交互工具用于操控视觉上下文、追踪路径和验证空间关系。为训练PERIA,我们设计了统一方案,结合监督式工具使用轨迹生成、复合奖励机制以及观察松弛组内组策略优化(OR-GIGPO)。在8个数据集的13个基准测试中,PERIA-8B相较于Qwen3-8B基线,在分布内任务上提升10.0%,分布外任务上提升4.4%;优于同规模先前最优模型7.0%–14.8%。其性能接近更大模型如Qwen3-VL-235B-A22B-Thinking和GPT-5,证明了工具增强对空间推理的有效性。
原文摘要 · Abstract (English)
While recent vision-language models (VLMs) demonstrate strong multimodal understanding, they remain limited in spatial reasoning tasks that require active evidence acquisition and multi-step visual interaction. This limitation suggests that relying solely on implicit visual representations from vision encoders is insufficient for recovering fine-grained spatial evidence. We introduce PERception-Interaction-reason Agent (PERIA), a tool-augmented visual agent for spatial reasoning tasks across map reasoning, visual probing, and vision reconstruction. PERIA uses two lightweight tool families: vision perception tools for exposing textual, symbolic, and spatial evidence, and vision interaction tools for manipulating visual context, tracing paths, and verifying spatial relations. To train PERIA, we develop a unified recipe that combines supervised tool-use trajectory synthesis, composite rewards, and Observation-Relaxed Group-in-Group Policy Optimization (OR-GIGPO) for effective multi-tool behavior. Experiments on 13 benchmarks from 8 datasets show that PERIA-8B improves over the Qwen3-8B backbone by 10.0% on in-distribution benchmarks and 4.4% on out-of-distribution benchmarks, while outperforming previous state-of-the-art baselines of similar size by 7.0%-14.8%. It also achieves performance comparable to much larger models such as Qwen3-VL-235B-A22B-Thinking and GPT-5, demonstrating the effectiveness of PERIA in enhancing spatial reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。