arXiv:2603.16016cs.CVcs.AI2026-03

从单张视角生成完整室内地板图,提升导航地图精度。

FlatLands: Generative Floormap Completion From a Single Egocentric View

  • 基于单视角图像生成完整鸟瞰地板图,端到端实现
  • 涵盖27万+真实场景数据,支持分布内与分布外评估
  • 适合研究具不确定性感知的智能体导航与生成式建图

单个第一人称视角图像通常只覆盖地面的一小部分,但完整的度量可达性地图能更好支持室内导航等应用。我们提出FlatLands,一个用于单视图鸟瞰(BEV)地板补全的数据集与基准。该数据集包含来自六个现有数据集的17,656个真实度量室内场景,共270,575个观测样本,提供对齐的观测、可见性、有效性及真值BEV地图,并包含分布内与分布外评估协议。我们对比了无需训练的方法、确定性模型、集成方法与随机生成模型。最终,我们将该任务实例化为一个端到端的单目RGB至地板图管道。FlatLands为不确定感知的室内映射与生成式补全提供了严格测试平台。

原文摘要 · Abstract (English)

A single egocentric image typically captures only a small portion of the floor, yet a complete metric traversability map of the surroundings would better serve applications such as indoor navigation. We introduce FlatLands, a dataset and benchmark for single-view bird's-eye view (BEV) floor completion. The dataset contains 270,575 observations from 17,656 real metric indoor scenes drawn from six existing datasets, with aligned observation, visibility, validity, and ground-truth BEV maps, and the benchmark includes both in- and out-of-distribution evaluation protocols. We compare training-free approaches, deterministic models, ensembles, and stochastic generative models. Finally, we instantiate the task as an end-to-end monocular RGB-to-floormaps pipeline. FlatLands provides a rigorous testbed for uncertainty-aware indoor mapping and generative completion for embodied navigation.

地板补全室内导航生成建图单目视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。