用深度信息提升视觉模型对空间结构的理解能力
DINOcular: Self-Supervised Visuospatial Representations

- 融合图像与深度图的块间和块内特征,学习联合视觉空间表征
- 在多个3D几何基准上超越同等规模模型,同时保持语义分割性能
- 适合需要精准空间感知的机器人、自动驾驶等场景
我们提出一种自监督框架,从RGB-D观测中学习联合视觉空间表征。尽管现代视觉基础模型几乎仅基于RGB图像训练,但许多具身系统可获取深度信息,提供单目图像无法恢复的几何细节。我们的方法通过块间与块内融合,将深度导出的几何先验与视觉主干网络结合,使模型高效编码外观与空间结构。所得表征在3D感知任务中表现优异,同时保持语义迁移能力:在多个3D几何基准上优于同等规模的现有方法,并在标准RGB-D语义分割任务中仍具竞争力。
原文摘要 · Abstract (English)
We introduce a self-supervised framework for learning joint visuospatial representations from RGB-D observations. While modern vision foundation models are trained almost exclusively on RGB images, many embodied systems have access to explicit depth sensing, which provides geometric information that monocular inputs cannot recover. Our method integrates depth-derived geometric priors with a visual backbone through inter-patch and intra-patch fusion, enabling the model to encode both appearance and spatial structure efficiently. The resulting representation shows promising improvements on 3D awareness while preserving semantic transfer: it outperforms prior methods of comparable scale on multiple 3D geometry benchmarks, and remains competitive when probed for standard RGB-D semantic segmentation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。