让视觉语言模型显式处理空间推理,提升导航任务准确率
Allocentric Perceiver: Disentangling Allocentric Reasoning from Egocentric Visual Priors via Frame Instantiation
- 用几何专家重建3D结构,构建与指令语义对齐的目标中心视角
- 在多类模型上实现10%的分配空间任务准确率提升
- 无需微调,适合需要精准空间理解的应用场景
随着视觉语言导航等空间任务需求上升,视觉语言模型(VLM)的分配空间感知能力受到关注。然而,现有VLM在需显式视角转换的分配空间查询中仍表现脆弱,其答案依赖目标中心视角而非观测相机视角。为此,我们提出Allocentric Perceiver,一种无需训练的方法:通过现成几何专家从一张或多张图像中恢复度量3D状态,并构建与查询条件匹配的分配参考系。通过将重建几何确定性地转换至目标坐标系,并以结构化、几何锚定的表示提示主干VLM,该方法将原本隐式的心理旋转转化为显式计算。我们在多个主干模型上评估该方法,在空间推理基准上一致获得约10%的分配任务性能提升,同时保持强姿态感知性能,超越经过空间感知微调的模型及当前开源与专有模型。
原文摘要 · Abstract (English)
With the rising need for spatially grounded tasks such as Vision-Language Navigation/Action, allocentric perception capabilities in Vision-Language Models (VLMs) are receiving growing focus. However, VLMs remain brittle on allocentric spatial queries that require explicit perspective shifts, where the answer depends on reasoning in a target-centric frame rather than the observed camera view. Thus, we introduce Allocentric Perceiver, a training-free strategy that recovers metric 3D states from one or more images with off-the-shelf geometric experts, and then instantiates a query-conditioned allocentric reference frame aligned with the instruction's semantic intent. By deterministically transforming reconstructed geometry into the target frame and prompting the backbone VLM with structured, geometry-grounded representations, Allocentric Perceriver offloads mental rotation from implicit reasoning to explicit computation. We evaluate Allocentric Perciver across multiple backbone families on spatial reasoning benchmarks, observing consistent and substantial gains ($\sim$10%) on allocentric tasks while maintaining strong egocentric performance, and surpassing both spatial-perception-finetuned models and state-of-the-art open-source and proprietary models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。