提出双模式3D场景图生成,快速粗粒度+慢速细粒度切换,提升开放任务效率。
Seeing Fast and Slow: Bimodal 3D Scene Graphs for Open-set Tasks
- 默认快速生成粗粒度场景图,需时关键时切换至精细模式。
- 相比开源最先进方法提速显著,支持实时任务执行集成。
- 适合需要动态调整细节级别的机器人开放任务场景。
开放任务执行可因根据上下文和环境探索中的信息演变,在粗粒度与细粒度场景表示间无缝切换而显著受益。例如,初始阶段使用粗粒度表示通常已足够,仅在探测到可能含任务相关物体的区域时才启用更精细的表示。为此,本文提出BiMoSG——一种面向开放任务的双模态3D场景图生成方法。BiMoSG默认采用‘快’模式高效生成粗粒度3D场景图,并可在必要时切换至‘慢’模式,生成针对任务相关物体的更细粒度开放词汇3D场景图。实验表明,该方法在生成速度上显著优于开源最先进方案,使场景图生成过程可与任务执行集成,实现真正意义上的实时部署。
原文摘要 · Abstract (English)
Open-set task execution can significantly benefit from seamlessly switching between coarse and fine scene representations depending on the context and the evolving information as the robot explores the environment. For example, it is often sufficient to start with a coarse scene representation initially and only employ a finer, more granular scene representation when the robot encounters regions which are likely to contain the task relevant objects. Hence, in this work, we propose BiMoSG, a bimodal 3D scene graph generation approach for open-set tasks. BiMoSG employs a "fast" mode by default to efficiently generate a coarse 3D scene graph and can switch to a "slow" mode for generating a finer open vocabulary 3D scene graph of task relevant objects. We demonstrate that our proposed 3D scene graph generation approach is significantly faster than the open-source state-of-the-art approaches. This allows us to integrate the scene graph generation process with task execution for real-time deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。