让3D场景能被精准指代:通过分组高斯点实现跨视角实例识别。
GroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian Splatting

- 用紧凑实例嵌入将高斯点分组,实现跨视角一致的3D实例结构。
- 在多个数据集上实现92.3%的指代分割准确率,优于基线方法。
- 适合需要精确语言交互的机器人导航、虚拟助手等应用。
同时重建与理解3D环境对具身智能体至关重要。现有前馈语义3D高斯泼溅(3DGS)方法虽能从稀疏多视角图像构建语义场景表示,但缺乏显式的实例区分能力,仅支持类别或短语级语义查询。为此,我们提出GroupForward,一种实例分组的前馈高斯泼溅模型,可从稀疏、无姿态、未标定的多视角图像中重建几何、外观、实例结构与语义。不同于将高维语义特征附加至每个高斯点的方法,GroupForward学习紧凑的实例嵌入,将高斯点分组为跨视角一致的3D实例,将前馈语义3DGS重构为基于实例级别的语义聚合与传播。在此基础上,我们进一步提出参考场景推理框架(RSRF),构建实例分组的3D场景图,并为给定指代表达检索候选实例。视觉-语言模型随后基于结构化实例证据与多视角观测进行推理,从候选中定位被指代的实例。实验表明,该实例分组重建与推理框架在语义重建与指代推理任务上均具有效性。
原文摘要 · Abstract (English)
Simultaneously reconstructing and understanding 3D environments is essential for embodied agents. Toward this goal, feed-forward semantic 3D Gaussian Splatting (3DGS) efficiently constructs semantic scene representations from sparse multi-view observations. However, existing methods lack explicit instance discrimination and mainly support category- or phrase-based semantic queries. To this end, we propose GroupForward, an instance-grouped feed-forward Gaussian splatting model that reconstructs geometry, appearance, instance structure, and semantics from sparse, unposed, and uncalibrated multi-view images. Unlike existing methods that attach high-dimensional semantic features to each Gaussian, GroupForward learns compact instance embeddings that group Gaussians into cross-view consistent 3D instances, reformulating feed-forward semantic 3DGS from per-Gaussian semantic feature rendering to instance-level semantic aggregation and propagation. Building on these instance groups, we further propose a Referential Scene Reasoning Framework (RSRF) for complex 3D referring segmentation. RSRF constructs an instance-grouped 3D scene graph and retrieves candidate instances for a given referring expression. A vision-language model then reasons over structured instance evidence and multi-view observations to identify the referred instance among the candidates. RSRF thereby extends language interaction from simple semantic querying to complex referential scene reasoning. Experiments on semantic reconstruction and referential reasoning demonstrate the effectiveness of our instance-grouped reconstruction and reasoning framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。