arXiv:2607.05247cs.CV2026-07被引 3

用边界建模提升视觉模型的空间感知能力,让AI更懂物体形状和位置。

Vision Pretraining for Dense Spatial Perception

论文配图:Vision Pretraining for Dense Spatial Perception
图 1 · 摘自论文原文
  • 通过动态学习亚像素级边界,以边界作为掩码目标来训练视觉表示
  • 在深度补全任务上实现从LingBot-Depth 1.0到2.0的升级,显著提升精度
  • 适合需要精细空间理解的机器人、自动驾驶等场景

密集空间感知是物理智能的核心,要求视觉系统从像素中恢复结构化、度量性且可操作的表征。当前视觉基础模型多侧重语义不变性,牺牲了细节空间理解。本文提出以边界为中心的视觉预训练方法,基于边界与形状不连续性对几何属性的重要提示作用,设计了掩码边界建模机制:动态学习亚像素级边界表示,并利用这些边界承载的标记作为掩码目标,促进稠密视觉标记的学习。通过扩展该框架,构建了LingBot-Vision模型,在多个下游视觉任务中表现优于DINOv3基线。值得注意的是,该模型推动了深度补全从LingBot-Depth 1.0到2.0的演进,显著提升了深度估计性能,成为具身人工智能的关键支撑。研究发现,边界建模不仅是简单线条识别,更是一种可扩展的预训练范式,能有效学习具有空间结构的视觉表征。

原文摘要 · Abstract (English)

Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observations. Modern visual foundation models tend to prioritize semantic invariance, often at the expense of detailed spatial understanding. In this work, we study vision pretraining through a boundary-centric lens, motivated by the premise that boundaries and shape discontinuities offer essential cues for perceiving geometric properties. Concretely, we propose masked boundary modeling, a self-supervised paradigm that dynamically learns sub-pixel boundary representations and subsequently leverages the discovered boundary-bearing tokens as masked targets to facilitate dense visual token learning. By scaling this framework, we develop LingBot-Vision and demonstrate its efficacy across a diverse set of downstream vision tasks with DINOv3 as a strong baseline. Remarkably, LingBot-Vision drives the progression from LingBot-Depth 1.0 to LingBot-Depth 2.0 for depth completion, and thereby yields enhanced depth estimation, a key pillar for embodied artificial intelligence. Our findings reveal that boundary modeling goes beyond simple line segments and instead serves as a scalable pretraining principle for learning spatially structured visual representations.

视觉预训练空间感知边界建模深度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。