arXiv:2509.20839cs.RO2025-09

让机器人提前猜出未知区域的结构和功能,提升导航效率。

SemSight: Probabilistic Bird's-Eye-View Prediction of Multi-Level Scene Semantics for Navigation

  • 用鸟瞰视角联合预测房间结构、场景上下文和目标分布
  • 在未探索区域预测准确率提升,结构一致性指标更优
  • 适合需要高效探索的自主导航系统使用

在目标驱动的导航与自主探索中,对未知区域的合理预测对高效导航和环境理解至关重要。现有方法多关注单个物体或几何占据图,缺乏建模房间级语义结构的能力。我们提出 SemSight,一种用于多层级场景语义的的概率鸟瞰视角预测模型。该模型联合推断结构布局、全局场景上下文及目标区域分布,在补全未探索区域语义地图的同时,估计目标类别的概率图。为训练 SemSight,我们在2000个室内布局图上模拟前沿驱动探索,构建了一个包含40,000组前后视点序列与完整语义地图的多样化数据集。采用编码器-解码器架构,并引入掩码约束监督策略:通过未探索区域的二值掩码,使监督仅聚焦于未知区域,迫使模型从已观察上下文中推断语义结构。实验结果表明,SemSight在未探索区域的关键功能类别预测性能显著提升,且在结构一致性(SC)和区域识别准确率(PA)等指标上优于非掩码监督方法。同时,在闭环模拟中提升了导航效率,减少了引导机器人到达目标区域所需的搜索步数。

原文摘要 · Abstract (English)

In target-driven navigation and autonomous exploration, reasonable prediction of unknown regions is crucial for efficient navigation and environment understanding. Existing methods mostly focus on single objects or geometric occupancy maps, lacking the ability to model room-level semantic structures. We propose SemSight, a probabilistic bird's-eye-view prediction model for multi-level scene semantics. The model jointly infers structural layouts, global scene context, and target area distributions, completing semantic maps of unexplored areas while estimating probability maps for target categories. To train SemSight, we simulate frontier-driven exploration on 2,000 indoor layout graphs, constructing a diverse dataset of 40,000 sequential egocentric observations paired with complete semantic maps. We adopt an encoder-decoder network as the core architecture and introduce a mask-constrained supervision strategy. This strategy applies a binary mask of unexplored areas so that supervision focuses only on unknown regions, forcing the model to infer semantic structures from the observed context. Experimental results show that SemSight improves prediction performance for key functional categories in unexplored regions and outperforms non-mask-supervised approaches on metrics such as Structural Consistency (SC) and Region Recognition Accuracy (PA). It also enhances navigation efficiency in closed-loop simulations, reducing the number of search steps when guiding robots toward target areas.

语义预测导航系统鸟瞰图多层级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。