让3D场景图边探索边理解,实时生成可查询的语义信息。
Think While You Map: Asynchronous Vision-Language Agents for Incremental 3D Scene Graphs

- 轻量级映射与重型语义模型并行运行,实现异步更新。
- 在三个视觉定位基准上提升15.3至18.8个[email protected],优于现有方法。
- 适合需要实时语义增强的机器人导航与交互系统。
开放词汇3D场景图方法通常分两阶段:先重建,再用视觉语言模型丰富语义,导致探索过程中图不可查询。我们提出一种异步架构,使轻量级在线映射与重型语义精炼并行。基于概率体素的骨干网络逐步保持对象身份稳定,背景VLM代理持续丰富图结构。该框架通过语义回环闭合解决重复追踪问题,附加细粒度视觉属性,并推导物体间空间关系。多目标帧调度器通过选择少量覆盖多个目标的信息帧,分摊VLM计算成本。最终场景图在探索过程中即可查询,且语义内容随时间不断丰富。本方法在语义分割任务(ScanNet、Replica)上达到或超越现有水平,在三个视觉定位基准(Sr3D+、Nr3D、ScanRefer)上相比先前最优结果提升15.3至18.8 [email protected]。
原文摘要 · Abstract (English)
Open-vocabulary 3D scene graph methods typically operate in two stages: first reconstruct, then enrich with vision-language models, leaving the graph unqueryable during exploration. We argue that this sequential coupling is unnecessary and propose an asynchronous architecture in which lightweight online mapping runs concurrently with heavyweight semantic refinement. A probabilistic voxel-based backbone maintains stable object identities incrementally, while background VLM agents progressively enrich the graph. This framework resolves duplicate object tracks through semantic loop closure, attaches fine-grained visual attributes and derives spatial relations between objects. A multi-target frame scheduler amortizes VLM cost by selecting a small set of informative frames that jointly cover multiple targets. The resulting scene graph is queryable during exploration and grows in semantic richness over time. Our method matches or outperforms existing open-vocabulary 3D scene graph methods on semantic segmentation (ScanNet, Replica) and surpasses the prior state-of-the-art across three visual grounding benchmarks (Sr3D+, Nr3D, ScanRefer) by 15.3 to 18.8 [email protected]. Project page: https://denizbickici.github.io/thinkgraphs/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。