arXiv:2604.19741cs.CV2026-04被引 1

用真实地理数据生成可导航的逼真城市视频

CityRAG: Stepping Into a City via Spatially-Grounded Video Generation

  • 基于地理注册数据,让视频生成锚定真实场景
  • 生成数分钟连续视频,保持天气光照一致
  • 适合自动驾驶与机器人仿真场景

我们解决生成三维一致、可导航且空间锚定的真实地点模拟问题。现有视频生成模型可依据文本(T2V)或图像(I2V)提示生成合理序列,但重建真实世界在任意天气与动态物体配置下的能力对自动驾驶和机器人仿真至关重要。为此,我们提出 CityRAG,一种利用大规模地理注册数据作为上下文,将生成结果锚定至物理场景,同时保留对复杂运动与外观变化的先验知识的视频生成模型。CityRAG 依赖时间未对齐的训练数据,使模型能语义解耦底层场景与其瞬时属性。实验表明,CityRAG 可生成连贯的数分钟长、物理合理的视频序列,在数千帧中保持天气与光照一致性,实现闭环,并沿复杂轨迹导航以重建真实地理结构。

原文摘要 · Abstract (English)

We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video generative models can produce a plausible sequence that is consistent with a text (T2V) or image (I2V) prompt. However, the capability to reconstruct the real world under arbitrary weather conditions and dynamic object configurations is essential for downstream applications including autonomous driving and robotics simulation. To this end, we present CityRAG, a video generative model that leverages large corpora of geo-registered data as context to ground generation to the physical scene, while maintaining learned priors for complex motion and appearance changes. CityRAG relies on temporally unaligned training data, which teaches the model to semantically disentangle the underlying scene from its transient attributes. Our experiments demonstrate that CityRAG can generate coherent minutes-long, physically grounded video sequences, maintain weather and lighting conditions over thousands of frames, achieve loop closure, and navigate complex trajectories to reconstruct real-world geography.

视频生成城市模拟空间锚定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。