用语言嵌入3D高斯点云实现开放词汇的自我中心场景理解
EgoSplat: Open-Vocabulary Egocentric Scene Understanding with Language Embedded 3D Gaussian Splatting
- 融合SAM2多视角特征,动态聚合实例语义信息
- 提升时空一致性,降低动态物体带来的重建伪影
- 在两个数据集上刷新性能,适合真实世界视觉应用
自我中心场景相比常规场景存在频繁遮挡、视角多样和动态交互等问题。遮挡与视角差异易引发多视图语义不一致,动态物体则可能作为瞬时干扰项,导致语义特征建模出现伪影。为此,本文提出EgoSplat,一种用于开放词汇自我中心场景理解的语言嵌入式3D高斯点云框架。设计多视图一致的实例特征聚合方法,利用SAM2的分割与追踪能力,选择性地跨视角聚合互补特征以精准表达每个实例的语义。同时构建实例感知的时空瞬态预测模块,通过融合多视角实例间的时空关联,提升预测的空间完整性与时间连续性,有效减少自我中心场景语义重建中的伪影。EgoSplat在两个数据集上均取得最先进性能:在ADT数据集上,定位准确率提升8.2%,分割mIoU提升3.7%,建立了开放词汇自我中心场景理解的新基准。代码将公开。
原文摘要 · Abstract (English)
Egocentric scenes exhibit frequent occlusions, varied viewpoints, and dynamic interactions compared to typical scene understanding tasks. Occlusions and varied viewpoints can lead to multi-view semantic inconsistencies, while dynamic objects may act as transient distractors, introducing artifacts into semantic feature modeling. To address these challenges, we propose EgoSplat, a language-embedded 3D Gaussian Splatting framework for open-vocabulary egocentric scene understanding. A multi-view consistent instance feature aggregation method is designed to leverage the segmentation and tracking capabilities of SAM2 to selectively aggregate complementary features across views for each instance, ensuring precise semantic representation of scenes. Additionally, an instance-aware spatial-temporal transient prediction module is constructed to improve spatial integrity and temporal continuity in predictions by incorporating spatial-temporal associations across multi-view instances, effectively reducing artifacts in the semantic reconstruction of egocentric scenes. EgoSplat achieves state-of-the-art performance in both localization and segmentation tasks on two datasets, outperforming existing methods with a 8.2% improvement in localization accuracy and a 3.7% improvement in segmentation mIoU on the ADT dataset, and setting a new benchmark in open-vocabulary egocentric scene understanding. The code will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。