GAIA-2可生成多视角可控驾驶视频,支持真实道路环境模拟。
GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
- 基于潜在扩散模型,融合车辆动态与环境语义进行可控生成
- 在英美德多地生成高分辨率、时空一致的多摄像头视频序列
- 适合自动驾驶系统开发中的复杂与罕见场景仿真
生成模型为模拟复杂环境提供了可扩展且灵活的范式,但现有方法难以满足自动驾驶的特定需求——如多智能体交互、细粒度控制和多摄像机一致性。我们提出GAIA-2(Generative AI for Autonomy),一种统一上述能力的潜在扩散世界模型。该模型支持基于多种结构化输入的可控视频生成:包括自车动力学、交通参与者配置、环境因素及道路语义。它能在地理分布多样的驾驶环境中生成高分辨率、时空一致的多摄像头视频(英国、美国、德国)。通过整合结构化条件与外部潜在嵌入(如来自专有驾驶模型),实现灵活且语义明确的场景合成。由此,GAIA-2可规模化生成常见与罕见驾驶场景,推动生成世界模型成为自动驾驶系统开发的核心工具。视频展示见https://wayve.ai/thinking/gaia-2。
原文摘要 · Abstract (English)
Generative models offer a scalable and flexible paradigm for simulating complex environments, yet current approaches fall short in addressing the domain-specific requirements of autonomous driving - such as multi-agent interactions, fine-grained control, and multi-camera consistency. We introduce GAIA-2, Generative AI for Autonomy, a latent diffusion world model that unifies these capabilities within a single generative framework. GAIA-2 supports controllable video generation conditioned on a rich set of structured inputs: ego-vehicle dynamics, agent configurations, environmental factors, and road semantics. It generates high-resolution, spatiotemporally consistent multi-camera videos across geographically diverse driving environments (UK, US, Germany). The model integrates both structured conditioning and external latent embeddings (e.g., from a proprietary driving model) to facilitate flexible and semantically grounded scene synthesis. Through this integration, GAIA-2 enables scalable simulation of both common and rare driving scenarios, advancing the use of generative world models as a core tool in the development of autonomous systems. Videos are available at https://wayve.ai/thinking/gaia-2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。