用真实首尔街景训练的视频世界模型,能生成百米级长视频。
Grounding World Simulation Models in a Real-World Metropolis
- 通过检索真实街景图像,引导自回归视频生成。
- 在首尔、釜山、安阿伯三城测试中表现最优,支持长距离轨迹。
- 适合城市规划、自动驾驶等需要真实环境模拟的场景。
如果世界模拟模型能渲染真实存在的城市而非虚构环境会怎样?现有生成式世界模型通过想象构建视觉上可信但人为的环境。我们提出首尔世界模型(SWM),一个基于真实首尔的城市级世界模型。SWM通过邻近街景图像的检索增强条件来锚定自回归视频生成。然而,该设计引入了若干挑战:检索参考与动态目标场景之间的时间错位、轨迹多样性受限及车载采集稀疏导致的数据稀疏性。我们通过跨时间配对、大规模合成数据集实现多样化相机轨迹,并设计视图插值流水线,从稀疏街景图像合成连贯训练视频。此外,我们引入虚拟前瞻汇点,通过持续将每个生成片段重新锚定至未来位置的检索图像来稳定长时程生成。我们在首尔、釜山和安阿伯三个城市评估SWM,结果表明其在生成空间准确、时间一致、长达数百米的真实城市环境视频方面优于现有方法,同时支持多样相机运动与文本提示的场景变化。
原文摘要 · Abstract (English)
What if a world simulation model could render not an imagined environment but a city that actually exists? Prior generative world models synthesize visually plausible yet artificial environments by imagining all content. We present Seoul World Model (SWM), a city-scale world model grounded in the real city of Seoul. SWM anchors autoregressive video generation through retrieval-augmented conditioning on nearby street-view images. However, this design introduces several challenges, including temporal misalignment between retrieved references and the dynamic target scene, limited trajectory diversity and data sparsity from vehicle-mounted captures at sparse intervals. We address these challenges through cross-temporal pairing, a large-scale synthetic dataset enabling diverse camera trajectories, and a view interpolation pipeline that synthesizes coherent training videos from sparse street-view images. We further introduce a Virtual Lookahead Sink to stabilize long-horizon generation by continuously re-grounding each chunk to a retrieved image at a future location. We evaluate SWM against recent video world models across three cities: Seoul, Busan, and Ann Arbor. SWM outperforms existing methods in generating spatially faithful, temporally consistent, long-horizon videos grounded in actual urban environments over trajectories reaching hundreds of meters, while supporting diverse camera movements and text-prompted scenario variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。