arXiv:2603.03283cs.CV2026-03被引 15

一个模型通吃各类点云数据,实现跨域统一表征。

Utonia: Toward One Encoder for All Point Clouds

  • 用自监督方法训练单一点云编码器,覆盖遥感、激光雷达、室内深度图等多源数据。
  • 在多领域联合训练下,模型学习到一致的表示空间并提升跨域感知能力。
  • 适用于机器人操控、视觉-语言-动作协同等多模态任务,适合3D基础模型研究者。

我们构想未来所有领域的点云数据能汇聚于单一模型,共同受益。为此,本文提出Utonia,首个在多种点云数据域(包括遥感、室外激光雷达、室内RGB-D序列、以物体为中心的CAD模型及仅由RGB视频生成的点云)上训练的自监督点变换器编码器。尽管这些数据在传感几何、密度和先验知识上差异显著,Utonia仍能学习到一致的表示空间,并实现跨域迁移。该统一架构不仅提升了感知性能,还揭示了仅在多域联合训练时才会出现的有趣涌现行为。此外,基于Utonia的特征可支持具身智能与多模态推理:将其用于视觉-语言-动作策略可改进机器人操作;融入视觉-语言模型后,在空间推理任务中也取得提升。我们希望Utonia能成为稀疏3D数据基础模型的一步,推动增强现实/虚拟现实、机器人与自动驾驶等下游应用发展。

原文摘要 · Abstract (English)

We dream of a future where point clouds from all domains can come together to shape a single model that benefits them all. Toward this goal, we present Utonia, a first step toward training a single self-supervised point transformer encoder across diverse domains, spanning remote sensing, outdoor LiDAR, indoor RGB-D sequences, object-centric CAD models, and point clouds lifted from RGB-only videos. Despite their distinct sensing geometries, densities, and priors, Utonia learns a consistent representation space that transfers across domains. This unification improves perception capability while revealing intriguing emergent behaviors that arise only when domains are trained jointly. Beyond perception, we observe that Utonia representations can also benefit embodied and multimodal reasoning: conditioning vision-language-action policies on Utonia features improves robotic manipulation, and integrating them into vision-language models yields gains on spatial reasoning. We hope Utonia can serve as a step toward foundation models for sparse 3D data, and support downstream applications in AR/VR, robotics, and autonomous driving.

点云编码自监督跨域表征3D基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。