arXiv:2606.03609cs.ROcs.LG2026-06

用3D可视域建模城市可通行空间,发现跨城市共性几何特征。

A 3D Isovist World Model -- Revealing a City's Unseen Geometry and Its Emergent Cross-City Signature

论文配图:A 3D Isovist World Model -- Revealing a City's Unseen Geometry and Its Emergent Cross-City Signature
图 1 · 摘自论文原文
  • 以3D可视域捕捉建筑间负空间,预测移动后视野变化。
  • 模型在曼哈顿与巴黎训练后,能识别城市级空间签名。
  • 适合做具身智能、机器人导航和城市分析的几何基础。

具身智能体导航城市依赖世界模型预测环境变化,但现有模型多关注外观而非实际可通行空间。多数几何模型将三维环境投影至平面,丢失了地上与多层结构。本文提出一种新方法:以3D可视域(isovist)建模建筑间的开放空间,即从某点向各方向测量到最近表面的距离。我们构建了一个具身世界模型,基于历史可视域和动作预测下一帧可视域,采用深度残差形式保持边缘清晰,并通过自回放调度采样确保上下文一致性,同时引入持久化的鸟瞰视角隐状态保证路径间连贯性。核心发现意外而深刻:一个不区分城市的模型在曼哈顿与巴黎数据上训练后,其时间隐状态中涌现出可线性解码的城市身份特征,显著优于单帧基线,表明该签名存在于学习动态中而非外观。该表征轻量、可解释、可复现,为具身人工智能、机器人和城市分析提供几何基础,配套开源数据集与流程。

原文摘要 · Abstract (English)

Embodied agents that navigate cities rely on world models that predict how their surroundings will change as they move. But for navigation, what matters is not what the buildings look like; it is where the agent can go. Most world models nonetheless predict appearance, learning how a scene looks rather than the space an agent can move through. Those that do target geometry, such as bird's-eye-view occupancy grids, flatten the three-dimensional environment onto a ground plane, discarding the above-ground and multi-level structure that shapes real navigation. What is missing is a predictive target that captures the navigable geometry an agent actually traverses, without photometric entanglement and without collapsing the third dimension. Our key idea is to model the open volume between buildings, the negative space, encoded as a 3D isovist: a spherical visibility-depth map recording the distance to the nearest surface in every direction. We introduce an embodied world model that predicts the next isovist from a short history of past isovists and a movement action. The prediction is formulated as a depth residual so the decoder inherits sharp building edges, trained with self-rollout scheduled sampling to keep corrupted context on the geometry manifold, and equipped with a persistent latent bird's-eye-view spatial map for cross-path consistency. Our central finding is emergent and unexpected: a single city-blind model trained on Manhattan and Paris develops a cross-city spatial signature, with city identity linearly decodable from its temporal latents far above single-frame baselines, so the signature lives in the learned dynamics rather than in appearance. The representation is lightweight, interpretable, and reproducible, offering a geometric substrate for spatial reasoning in embodied AI, robotics, and urban analysis, released with an open dataset and pipeline.

具身智能3D建模城市分析空间表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。