arXiv:2604.21190cs.CV2026-04

用动态协作的多模型系统提升空间推理能力

SpatiO: Adaptive Test-Time Orchestration of Vision-Language Agents for Spatial Reasoning

论文配图:SpatiO: Adaptive Test-Time Orchestration of Vision-Language Agents for Spatial Reasoning
图 1 · 摘自论文原文
  • 引入异构视觉语言代理,各具不同先验知识
  • 推理时动态评估并重加权代理,提升可靠性
  • 在多个基准上超越主流方法,适合复杂场景理解

理解视觉场景不仅需要识别物体,还需推理其空间关系。空间推理需融合2D外观、深度信号和几何约束等多种归纳偏置,而这些偏置的可靠性随上下文变化。因此,有效空间推理需具备空间适应性:根据输入灵活协调不同推理策略。但现有方法多依赖单一推理流程,隐含固定空间先验,难以应对分布变化。多代理系统可通过聚合多样推理路径提供替代方案,但此前研究多采用同质代理,限制了归纳偏置多样性。本文提出SpatiO,一种异构多代理框架,协调多个具有互补归纳偏置的视觉语言专家。为实现高效协作,我们提出测试时编排(TTO)机制,在不修改模型参数的前提下,动态评估并重加权代理的可靠性。在3DSRBench、STVQA-7k、CV-Bench和Omni3D-Bench等多个空间推理基准上的实验表明,SpatiO持续优于闭源与开源基线。

原文摘要 · Abstract (English)

Understanding visual scenes requires not only recognizing objects but also reasoning about their spatial relationships. Unlike general vision-language tasks, spatial reasoning requires integrating multiple inductive biases, such as 2D appearance cues, depth signals, and geometric constraints, whose reliability varies across contexts. This suggests that effective spatial reasoning requires spatial adaptability: the ability to flexibly coordinate different reasoning strategies depending on the input. However, most existing approaches rely on a single reasoning pipeline that implicitly learns a fixed spatial prior, limiting their ability to adapt under distribution changes. Multi-agent systems offer a promising alternative by aggregating diverse reasoning trajectories, but prior attempts in spatial reasoning primarily employ homogeneous agents, restricting the diversity of inductive biases they can leverage. In this work, we introduce SpatiO, a heterogeneous multi-agent framework for spatial reasoning that coordinates multiple vision-language specialists with complementary inductive biases. To enable effective collaboration, we propose Test-Time Orchestration (TTO), an calibration mechanism that dynamically evaluates and reweights agents based on their observed reliability during inference, without modifying model parameters. Extensive experiments on diverse spatial reasoning benchmarks, including 3DSRBench, STVQA-7k, CV-Bench, and Omni3D-Bench, demonstrate that SpatiO consistently improves spatial reasoning performance over both closed-source and open-source baselines. The project page is available at https://cy-h1329.github.io/spatio/.

空间推理多智能体视觉语言自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。