arXiv:2504.06994cs.ROcs.AI2025-04被引 37

提出统一表示法,让机器人高效感知近距与远距语义信息。

RayFronts: Open-Set Semantic Ray Frontiers for Online Scene Understanding and Exploration

  • 用体素和射线联合编码近远距离语义,支持开放集映射。
  • 零样本3D语义分割性能提升1.34倍,吞吐量提高16.5倍。
  • 适用于实时导航与探索,适合高动态开放世界任务。

开放集语义映射对开放世界的机器人至关重要。现有方法或受限于深度范围,或仅在受限环境下映射远距离物体,难以融合近距与远距观测,且在细粒度语义与效率间存在权衡。本文提出 RayFronts,一种统一表示,实现密集与远距离的高效语义建图。该方法将任务无关的开放集语义编码至近距体素与远距边界射线中,显著减少搜索空间,使机器人在感知范围内与范围外均能做出明智决策,且在 Orin AGX 上运行速度达 8.84 Hz。基准测试显示,其近距语义表现优于基线:零样本3D语义分割性能提升1.34倍,吞吐量提高16.5倍。传统在线建图评估常受其他系统组件干扰,本文提出无规划器依赖的评估框架,有效衡量远距搜索与探索能力,结果显示,RayFronts 比最接近的在线基线减少搜索体积2.2倍。

原文摘要 · Abstract (English)

Open-set semantic mapping is crucial for open-world robots. Current mapping approaches either are limited by the depth range or only map beyond-range entities in constrained settings, where overall they fail to combine within-range and beyond-range observations. Furthermore, these methods make a trade-off between fine-grained semantics and efficiency. We introduce RayFronts, a unified representation that enables both dense and beyond-range efficient semantic mapping. RayFronts encodes task-agnostic open-set semantics to both in-range voxels and beyond-range rays encoded at map boundaries, empowering the robot to reduce search volumes significantly and make informed decisions both within & beyond sensory range, while running at 8.84 Hz on an Orin AGX. Benchmarking the within-range semantics shows that RayFronts's fine-grained image encoding provides 1.34x zero-shot 3D semantic segmentation performance while improving throughput by 16.5x. Traditionally, online mapping performance is entangled with other system components, complicating evaluation. We propose a planner-agnostic evaluation framework that captures the utility for online beyond-range search and exploration, and show RayFronts reduces search volume 2.2x more efficiently than the closest online baselines.

语义建图开放集实时感知机器人探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。