arXiv:2609.06004cs.CVcs.AI2026-09

让视觉语言模型在测试时自适应几何约束,提升空间推理准确性。

Geometry-Aware Test-Time Learning for Quantitative Spatial Reasoning

论文配图:Geometry-Aware Test-Time Learning for Quantitative Spatial Reasoning
图 1 · 摘自论文原文
  • 用几何一致性约束和测试数据生成伪标签,动态调整模型
  • 在Q-Spatial-ScanNet上使两个模型准确率分别提升6.47%和9.41%
  • 适合需要高精度空间推理的场景,如机器人导航与三维理解

视觉语言模型(VLM)中的定量空间推理旨在从2D图像和自然语言查询中推断3D空间内物体间的距离与方向关系。尽管近期取得进展,但在分布偏移下仍表现脆弱,主要因3D标注成本过高。模型常对新物体配置或重述查询产生不一致预测,暴露出表征与真实几何的错位。为此,我们提出TTL-SR:一种面向定量空间推理的几何感知测试时学习框架,利用几何一致性约束与无标注测试数据,适配目标域。具体而言,通过添加几何关联的辅助查询增强输入,以自适应几何触发过滤不可靠预测,构建结构化标记级别的伪标签,并仅使用测试数据在几何感知多目标损失下更新参数。实验表明,TTL-SR显著提升空间推理性能,在Q-Spatial-ScanNet数据集上使Qwen3-VL-4B-Instruct与SpatialRGPT-VILA-1.5-8B模型准确率分别提升6.47%和9.41%。

原文摘要 · Abstract (English)

Quantitative spatial reasoning in visual-language models (VLMs) aims to infer spatial distances and directional relationships among objects in 3D space from a 2D image and a natural language query. Despite recent progress, VLM spatial reasoning remains brittle under distribution shifts, largely due to the high cost of 3D supervision. As a result, models often produce inconsistent or contradictory predictions when faced with novel object configurations or rephrased spatial queries, revealing a misalignment between learned representations and underlying geometry. To address this, we propose TTL-SR, a geometry-aware Test-Time Learning framework for quantitative Spatial Reasoning that leverages geometric consistency constraints and unlabeled test data to adapt models to target domains. Specifically, TTL-SR augments the input query with geometrically coupled auxiliary queries, filters unreliable predictions via adaptive geometric triggering to construct structured token-level pseudo-labels, and updates model parameters under a geometry-aware multi-objective loss using only test data. Experimental results demonstrate that TTL-SR significantly boosts spatial reasoning performance, yielding 6.47% and 9.41% accuracy gains for Qwen3-VL-4B-Instruct and SpatialRGPT-VILA-1.5-8B on Q-Spatial-ScanNet dataset, respectively.

空间推理测试时学习几何感知视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。