arXiv:2510.15018cs.CVcs.AI2025-10中稿 · ICLR被引 6

用城市旅游视频自动生成高真实感城市仿真环境,提升智能体训练效果。

UrbanVerse: Scaling Urban Simulation by Watching City-Tour Videos

  • 从众包城市视频中自动提取场景布局并生成可交互仿真
  • 在160个真实城市场景上实现+6.3%仿真成功率和+30.1%零样本迁移表现
  • 适合城市机器人、自动驾驶等需要大规模真实场景训练的研究者

城市具身智能体(如配送机器人、四足机器人)正日益活跃于城市街道,需在复杂环境中完成最后一公里连接。现有手工或程序生成的仿真场景难以兼顾规模与真实复杂性。我们提出UrbanVerse,一种数据驱动的真实世界到仿真系统,将众包城市旅游视频转化为物理感知、可交互的仿真场景。UrbanVerse包含:(i) UrbanVerse-100K,一个超过10万条标注的城市3D资产库,带有语义与物理属性;(ii) UrbanVerse-Gen,一个自动管道,从视频中提取场景布局,并使用检索到的资产构建度量级3D仿真。在IsaacSim中,系统生成了来自24个国家的160个高质量场景,并提供10个艺术家设计的测试场景基准。实验表明,UrbanVerse场景保持真实语义与布局,人类评估的真实感媲美人工构建场景。在城市导航任务中,基于UrbanVerse训练的策略表现出缩放规律与强泛化能力,在仿真中成功率提升+6.3%,零样本仿真到现实迁移提升+30.1%,仅用两次干预即完成300米真实世界任务。

原文摘要 · Abstract (English)

Urban embodied AI agents, ranging from delivery robots to quadrupeds, are increasingly populating our cities, navigating chaotic streets to provide last-mile connectivity. Training such agents requires diverse, high-fidelity urban environments to scale, yet existing human-crafted or procedurally generated simulation scenes either lack scalability or fail to capture real-world complexity. We introduce UrbanVerse, a data-driven real-to-sim system that converts crowd-sourced city-tour videos into physics-aware, interactive simulation scenes. UrbanVerse consists of: (i) UrbanVerse-100K, a repository of 100k+ annotated urban 3D assets with semantic and physical attributes, and (ii) UrbanVerse-Gen, an automatic pipeline that extracts scene layouts from video and instantiates metric-scale 3D simulations using retrieved assets. Running in IsaacSim, UrbanVerse offers 160 high-quality constructed scenes from 24 countries, along with a curated benchmark of 10 artist-designed test scenes. Experiments show that UrbanVerse scenes preserve real-world semantics and layouts, achieving human-evaluated realism comparable to manually crafted scenes. In urban navigation, policies trained in UrbanVerse exhibit scaling power laws and strong generalization, improving success by +6.3% in simulation and +30.1% in zero-shot sim-to-real transfer comparing to prior methods, accomplishing a 300 m real-world mission with only two interventions.

城市仿真视频生成具身智能真实感建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。