从多视角图像直接生成对象级3D令牌组,实现端到端场景理解与编辑。
Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views

- 用实例令牌与锚点令牌的双层结构分解场景,分离身份与外观信息。
- 无需3D标注,通过可微渲染联合优化重建与分割,在实例分割上超越基线。
- 支持实例级编辑和开放词汇检索,适合需要交互式操作的3D应用。
三维场景的理解应基于其物体,而非构成它们的原始几何体。然而,当前前馈式重建方法输出密集且无结构的点或高斯分布,物体级别的结构需事后恢复。本文提出一种前馈框架,直接从无姿态的多视角图像中分解出实例结构化的3D令牌组——紧凑的对象中心单元,由此可直接进行重建、分割与操作。每个令牌组包含一个捕获实体身份的实例令牌和编码局部几何与外观的锚点令牌,解码为一组3D高斯分布。这种两层因子分解将物体身份与局部外观解耦,使物体实例成为表示的原生接口而非衍生产物。令牌组通过可微渲染联合重建与分割监督学习,无需3D标注。该前馈模型在类别无关实例分割上优于逐场景优化基线,同时在新视图合成上保持竞争力。此外,同一令牌组可直接用于实例级场景编辑(删除、移动、插入物体)及高效开放词汇3D实例检索,检索复杂度随实例数量而非原始几何体数量增长。
原文摘要 · Abstract (English)
A 3D scene is understood through its objects, not the primitives that compose them. Yet feed-forward reconstruction methods output dense, unstructured sets of points or Gaussians, leaving object-level structure to be recovered after the fact. We propose a feed-forward framework that decomposes a scene into instance-structured 3D token groups directly from unposed multi-view images -- compact object-centric units from which reconstruction, segmentation, and manipulation all follow. Each token group pairs an instance token capturing entity-level identity with anchor tokens that encode local geometry and appearance, which are decoded into a set of 3D Gaussians. This two-level factorization decouples object identity from local appearance, making object instances a native interface of the representation rather than a derived product. The token groups are learned through differentiable rendering with joint reconstruction and segmentation supervision, requiring no 3D annotations. Our feed-forward model surpasses per-scene optimization baselines in class-agnostic instance segmentation while remaining competitive in novel view synthesis. Beyond these metrics, the same token groups directly unlock instance-level scene editing -- removing, translating, or inserting objects by operating on their groups -- as well as efficient open-vocabulary 3D instance retrieval, where retrieval complexity scales with the number of instances rather than primitives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。