X-World可生成可控的多视角驾驶视频,支持长期稳定仿真。
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
- 基于多视角视频生成,以动作和环境控制信号驱动未来画面
- 长时序滚动保持动态稳定,多视角一致性高,动作跟随精准
- 适合用于自动驾驶算法的大规模可复现仿真测试
在端到端自动驾驶时代,视觉-语言-动作(VLA)策略直接将原始传感器数据映射为驾驶动作,但当前评估仍严重依赖真实道路测试,存在成本高、场景覆盖有限且难以复现的问题。为此,我们提出X-World——一种动作条件下的多相机生成式世界模型,直接在视频空间中模拟未来观测。给定同步多视角相机历史与未来动作序列,X-World生成符合指令动作的多视角视频流。为实现可复现与可编辑的场景推演,该模型还支持对动态交通参与者和静态道路元素的可选控制,并保留文本提示接口以调节外观(如天气、时段)。此外,通过外观提示可实现视频风格迁移,同时保持动作与场景动态不变。核心是多视角潜在视频生成器,显式增强跨视角几何一致性和时间连贯性。实验表明,X-World在(i)多视角一致性、(ii)长时序动态稳定性、(iii)高可控性(严格遵循动作与场景控制)方面表现优异,为大规模、可复现的评估提供了实用基础。
原文摘要 · Abstract (English)
Scalable and reliable evaluation is increasingly critical in the end-to-end era of autonomous driving, where vision--language--action (VLA) policies directly map raw sensor streams to driving actions. Yet, current evaluation pipelines still rely heavily on real-world road testing, which is costly, biased toward limited scenario coverage, and difficult to reproduce. These challenges motivate a real-world simulator that can generate realistic future observations under proposed actions, while remaining controllable and stable over long horizons. We present X-World, an action-conditioned multi-camera generative world model that simulates future observations directly in video space. Given synchronized multi-view camera history and a future action sequence, X-World generates future multi-camera video streams that follow the commanded actions. To ensure reproducible and editable scene rollouts, X-World further supports optional controls over dynamic traffic agents and static road elements, and retains a text-prompt interface for appearance-level control (e.g., weather and time of day). Beyond world simulation, X-World also enables video style transfer by conditioning on appearance prompts while preserving the underlying action and scene dynamics. At the core of X-World is a multi-view latent video generator designed to explicitly encourage cross-view geometric consistency and temporal coherence under diverse control signals. Experiments show that X-World achieves high-quality multi-view video generation with (i) strong view consistency across cameras, (ii) stable temporal dynamics over long rollouts, and (iii) high controllability with strict action following and faithful adherence to optional scene controls. These properties make X-World a practical foundation for scalable and reproducible evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。