让学习提出假设,几何决定结果,提升空间推理可靠性
Learning Proposes, Geometry Disposes: A Modular Framework for Efficient Spatial Reasoning
- 学习模块生成姿态和深度猜测,几何算法最终判定结果
- 在TUM数据集上,合理对齐后性能显著提升
- 适合需要鲁棒空间感知的机器人与AR应用
空间感知旨在从视觉观测中估计相机运动和场景结构,传统方法依赖几何建模与物理一致性约束。近年来学习方法在几何感知上展现出强大表征能力,常被用于增强经典几何系统。然而,学习组件是否应直接替代几何估计,或仅作为中间模块仍存疑问。本文提出端到端模块化框架,让学习提出几何假设,几何算法做出最终判断。以RGB-D序列上的相对相机位姿估计为例,使用VGGT生成姿态与深度提议,再通过经典点到平面ICP作为几何后端。在TUM RGB-D基准测试中发现:(1)仅靠学习提出姿态不可靠;(2)若学习提议未正确对齐相机内参,性能反而下降;(3)当学习提议的深度经几何对齐后,结合几何处置阶段,在中等挑战的刚体场景下表现持续提升。结果表明,几何不仅是修正环节,更是验证与吸收学习观测的关键仲裁者。研究强调了模块化、几何感知设计对鲁棒空间感知的重要性。
原文摘要 · Abstract (English)
Spatial perception aims to estimate camera motion and scene structure from visual observations, a problem traditionally addressed through geometric modeling and physical consistency constraints. Recent learning-based methods have demonstrated strong representational capacity for geometric perception and are increasingly used to augment classical geometry-centric systems in practice. However, whether learning components should directly replace geometric estimation or instead serve as intermediate modules within such pipelines remains an open question. In this work, we address this gap and investigate an end-to-end modular framework for effective spatial reasoning, where learning proposes geometric hypotheses, while geometric algorithms dispose estimation decisions. In particular, we study this principle in the context of relative camera pose estimation on RGB-D sequences. Using VGGT as a representative learning model, we evaluate learning-based pose and depth proposals under varying motion magnitudes and scene dynamics, followed by a classical point-to-plane RGB-D ICP as the geometric backend. Our experiments on the TUM RGB-D benchmark reveal three consistent findings: (1) learning-based pose proposals alone are unreliable; (2) learning-proposed geometry, when improperly aligned with camera intrinsics, can degrade performance; and (3) when learning-proposed depth is geometrically aligned and followed by a geometric disposal stage, consistent improvements emerge in moderately challenging rigid settings. These results demonstrate that geometry is not merely a refinement component, but an essential arbiter that validates and absorbs learning-based geometric observations. Our study highlights the importance of modular, geometry-aware system design for robust spatial perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。