让3D场景分割支持开放词汇,通过跨视角融合提升细节准确率
OpenInsGaussian: Open-vocabulary Instance Gaussian Segmentation with Context-aware Cross-view Fusion
- 为每个实例掩码注入丰富语义上下文,增强识别能力
- 通过注意力机制融合多视角特征,减少对齐误差和信息缺失
- 适用于自动驾驶等需要精准3D理解的现实场景
三维场景理解对自动驾驶、机器人和增强现实至关重要。现有基于语义高斯点云的方法利用大规模二维视觉模型将2D语义特征投影到3D场景,但存在两大缺陷:(1) 预处理阶段个体掩码缺乏充分上下文提示;(2) 多视角特征融合时出现不一致与细节丢失。本文提出 extbf{OpenInsGaussian},一个具备上下文感知跨视角融合能力的开放词汇实例高斯分割框架。方法包含两个模块:上下文感知特征提取,为每个掩码补充丰富语义上下文;注意力驱动特征聚合,选择性融合多视角特征以缓解对齐错误与不完整问题。在基准数据集上的大量实验表明,OpenInsGaussian 在开放词汇3D高斯分割任务中达到领先性能,显著超越现有基线。结果验证了该方法的鲁棒性与通用性,为三维场景理解及其在多样化真实场景中的应用迈出了重要一步。
原文摘要 · Abstract (English)
Understanding 3D scenes is pivotal for autonomous driving, robotics, and augmented reality. Recent semantic Gaussian Splatting approaches leverage large-scale 2D vision models to project 2D semantic features onto 3D scenes. However, they suffer from two major limitations: (1) insufficient contextual cues for individual masks during preprocessing and (2) inconsistencies and missing details when fusing multi-view features from these 2D models. In this paper, we introduce \textbf{OpenInsGaussian}, an \textbf{Open}-vocabulary \textbf{Ins}tance \textbf{Gaussian} segmentation framework with Context-aware Cross-view Fusion. Our method consists of two modules: Context-Aware Feature Extraction, which augments each mask with rich semantic context, and Attention-Driven Feature Aggregation, which selectively fuses multi-view features to mitigate alignment errors and incompleteness. Through extensive experiments on benchmark datasets, OpenInsGaussian achieves state-of-the-art results in open-vocabulary 3D Gaussian segmentation, outperforming existing baselines by a large margin. These findings underscore the robustness and generality of our proposed approach, marking a significant step forward in 3D scene understanding and its practical deployment across diverse real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。