arXiv:2504.19266cs.CV2025-04被引 4

实时开放词汇场景理解,提升3D感知精度与响应速度。

OpenFusion++: An Open-vocabulary Real-time Scene Understanding System

  • 融合基础模型置信图优化点云,动态更新语义标签。
  • 在ICL、Replica等数据集上显著提升语义准确率和查询响应速度。
  • 适合需要实时交互的视觉语言导航与增强现实应用。

实时开放词汇场景理解对于视觉-语言导航、具身智能和增强现实等应用中的高效3D感知至关重要。然而,现有方法存在实例分割不精确、语义更新静态以及复杂查询处理能力有限等问题。为此,我们提出OpenFusion++,一种基于TSDF的实时3D语义-几何重建系统。该方法通过融合基础模型生成的置信图来优化3D点云,利用基于实例区域的自适应缓存动态更新全局语义标签,并采用双路径编码框架,将物体属性与环境上下文结合,实现精准查询响应。在ICL、Replica、ScanNet和ScanNet++数据集上的实验表明,OpenFusion++在语义准确性和查询响应速度方面均显著优于基线方法。

原文摘要 · Abstract (English)

Real-time open-vocabulary scene understanding is essential for efficient 3D perception in applications such as vision-language navigation, embodied intelligence, and augmented reality. However, existing methods suffer from imprecise instance segmentation, static semantic updates, and limited handling of complex queries. To address these issues, we present OpenFusion++, a TSDF-based real-time 3D semantic-geometric reconstruction system. Our approach refines 3D point clouds by fusing confidence maps from foundational models, dynamically updates global semantic labels via an adaptive cache based on instance area, and employs a dual-path encoding framework that integrates object attributes with environmental context for precise query responses. Experiments on the ICL, Replica, ScanNet, and ScanNet++ datasets demonstrate that OpenFusion++ significantly outperforms the baseline in both semantic accuracy and query responsiveness.

3D感知开放词汇实时系统语义重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。