一次性完成3D重建与实例理解,无需逐场景优化
InstanceSplat: Instance-Aware Feed-Forward 3D Gaussian Splatting for Scene Understanding

- 单次前向传播构建含外观、几何、实例身份的统一高斯表示
- 在多视图下实现跨视角一致的实例特征,支持新视角合成与分割
- 适合需要高效泛化能力的3D视觉任务,如自动驾驶与机器人感知
前馈式3D高斯点云(3DGS)实现了高效的通用3D重建,但现有方法在场景理解上仍以类别为导向。而实例感知的3DGS通常依赖每场景优化,且将重建与实例和语义学习解耦,限制了三者间的互惠。我们提出InstanceSplat,一种统一的前馈式3DGS框架,可从无位姿的多视图图像中实现通用3D重建与实例感知的场景理解。单次前向传播中,InstanceSplat构建一个联合编码外观、几何、实例身份与语言对齐语义的实例感知高斯表示。共享的3D高斯体在不同视角间锚定实例身份,生成可渲染且跨视角一致的实例特征。为促进重建与场景理解相互增强,我们设计了以实例为中心的学习策略,通过共享实例结构连接重建、实例学习与语义学习。具体而言,实例线索引导重建,语言对齐语义强化同类混淆实例的区分能力,实例区域将语义证据聚合为连贯的对象级预测。在新视角合成、实例分割与开放词汇语义理解任务上,于不同输入视图设置及未见数据集上的实验均表明其达到领先性能,具备实际效率与强泛化能力。
原文摘要 · Abstract (English)
Feed-forward 3D Gaussian Splatting (3DGS) enables efficient and generalizable 3D reconstruction, but current feed-forward 3DGS methods for scene understanding remain largely category-oriented. In contrast, instance-aware 3DGS methods typically rely on per-scene optimization and often decouple reconstruction from instance and semantic learning, limiting reciprocal interactions among them. We present InstanceSplat, a unified feed-forward 3DGS framework for generalizable 3D reconstruction and instance-aware scene understanding from pose-free multi-view images. In a single forward pass, InstanceSplat constructs an instance-aware Gaussian representation that jointly encodes appearance, geometry, instance identity, and language-aligned semantics. Shared 3D Gaussians ground instance identities across views, producing renderable and cross-view-consistent instance features. To allow reconstruction and scene understanding to benefit from each other, we further design an instance-centric learning strategy that connects reconstruction, instance learning, and semantic learning through shared instance structure. Specifically, instance cues guide reconstruction, language-aligned semantics strengthen the discrimination of confusing same-category instances, and instance regions aggregate semantic evidence into coherent object-level predictions. Experiments on novel-view synthesis, instance segmentation, and open-vocabulary semantic understanding under varying input-view settings and on an unseen dataset demonstrate state-of-the-art performance, practical efficiency, and strong generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。