arXiv:2511.16567cs.CV2025-11被引 3

用点图实现3D场景理解,无需图像仅靠坐标就能学出强模型。

POMA-3D: The Point Map Way to 3D Scene Understanding

  • 通过点图编码3D坐标,兼容2D模型输入格式。
  • 在6.5K房间级场景上训练,支持多种3D任务性能领先。
  • 适合做3D视觉、机器人导航等只给几何信息的场景理解。

本文提出POMA-3D,首个从点图学习的自监督3D表示模型。点图将3D坐标显式编码在规则2D网格上,保留全局3D几何结构,同时兼容2D基础模型输入格式。为迁移丰富的2D先验知识,设计视图到场景对齐策略;由于点图依赖于规范空间视角,引入POMA-JEPA联合嵌入预测架构,强制多视角间几何一致的特征表达。此外,构建了ScenePoint数据集,包含6.5K个房间级RGB-D场景和100万张2D图像场景,支持大规模POMA-3D预训练。实验表明,POMA-3D作为强大骨干网络,在3D问答、具身导航、场景检索和具身定位等任务中表现优异,仅使用几何输入(即3D坐标)即可达成。整体上,POMA-3D探索了一条基于点图的3D场景理解新路径,缓解3D表征学习中预训练先验稀缺与数据不足的问题。

原文摘要 · Abstract (English)

In this paper, we introduce POMA-3D, the first self-supervised 3D representation model learned from point maps. Point maps encode explicit 3D coordinates on a structured 2D grid, preserving global 3D geometry while remaining compatible with the input format of 2D foundation models. To transfer rich 2D priors into POMA-3D, a view-to-scene alignment strategy is designed. Moreover, as point maps are view-dependent with respect to a canonical space, we introduce POMA-JEPA, a joint embedding-predictive architecture that enforces geometrically consistent point map features across multiple views. Additionally, we introduce ScenePoint, a point map dataset constructed from 6.5K room-level RGB-D scenes and 1M 2D image scenes to facilitate large-scale POMA-3D pretraining. Experiments show that POMA-3D serves as a strong backbone for both specialist and generalist 3D understanding. It benefits diverse tasks, including 3D question answering, embodied navigation, scene retrieval, and embodied localization, all achieved using only geometric inputs (i.e., 3D coordinates). Overall, our POMA-3D explores a point map way to 3D scene understanding, addressing the scarcity of pretrained priors and limited data in 3D representation learning. Project Page: https://matchlab-imperial.github.io/poma3d/

3D理解点图自监督具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。