arXiv:2604.09045cs.CV2026-04被引 1

让3D高斯点云自动学会跨场景的物体级表示,无需逐场景调参。

Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting

  • 用预训练的全局物体中心学习模块,构建跨场景一致的物体代码本。
  • 直接用无监督物体掩码监督3D高斯点,无需额外处理或对齐。
  • 提升机器人交互等下游任务的泛化能力,适合多场景应用。

现有3D场景理解方法利用视觉基础模型的2D掩码监督辐射场,实现实例级3D分割。但基础模型提供的监督信号并非本质物体中心,常需额外掩码预/后处理或专门训练与损失设计来解决多视角掩码身份冲突。所学物体身份依赖具体场景,限制了跨场景泛化。为此,我们提出一种数据集级别的物体中心监督方案,用于3D高斯喷溅(3DGS)中的物体表示学习。基于预训练的槽注意力式全局物体中心学习(GOCL)模块,我们学习一个场景无关的物体代码本,提供跨视图与跨场景的一致、身份锚定表示。通过将代码本与模块的无监督物体掩码结合,可直接监督3D高斯点的身份特征,无需额外掩码处理或显式多视图对齐。所学场景无关代码本使物体监督与识别无需每场景微调或重训练。本方法首次将无监督物体中心学习(OCL)引入3DGS,获得更结构化的表示,并在机器人交互、场景理解及跨场景泛化等下游任务中表现更优。

原文摘要 · Abstract (English)

Recent works on 3D scene understanding leverage 2D masks from visual foundation models (VFMs) to supervise radiance fields, enabling instance-level 3D segmentation. However, the supervision signals from foundation models are not fundamentally object-centric and often require additional mask pre/post-processing or specialized training and loss design to resolve mask identity conflicts across views. The learned identity of the 3D scene is scene-dependent, limiting generalizability across scenes. Therefore, we propose a dataset-level, object-centric supervision scheme to learn object representations in 3D Gaussian Splatting (3DGS). Building on a pre-trained slot attention-based Global Object Centric Learning (GOCL) module, we learn a scene-agnostic object codebook that provides consistent, identity-anchored representations across views and scenes. By coupling the codebook with the module's unsupervised object masks, we can directly supervise the identity features of 3D Gaussians without additional mask pre-/post-processing or explicit multi-view alignment. The learned scene-agnostic codebook enables object supervision and identification without per-scene fine-tuning or retraining. Our method thus introduces unsupervised object-centric learning (OCL) into 3DGS, yielding more structured representations and better generalization for downstream tasks such as robotic interaction, scene understanding, and cross-scene generalization.

3D高斯喷溅物体中心无监督学习跨场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。