让自动驾驶视频生成更真实,保持物体一致性与空间精准度。
InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation
- 引入实例流引导与空间对齐机制,提升视频时序与几何一致性。
- 在nuScenes数据集上实现当前最佳生成质量,支持下游驾驶任务。
- 可生成稀有高危场景,适合自动驾驶系统安全性评估。
自动驾驶依赖高质量、大规模多视角驾驶视频训练鲁棒模型。虽然世界模型能低成本生成逼真驾驶视频,但难以维持实例级时间一致性与空间几何精度。为此,我们提出InstaDrive,通过两项关键改进:(1) 实例流引导器,跨帧提取并传播实例特征,确保实例身份长期一致;(2) 空间几何对齐器,增强空间推理能力,精确控制实例位置,并显式建模遮挡层级。结合这些实例感知机制,InstaDrive在nuScenes数据集上达到当前最优视频生成质量,显著提升下游自动驾驶任务性能。此外,我们利用CARLA自动驾驶系统,在多样地图与区域中程序化、随机生成罕见但关键安全的驾驶场景,为自动驾驶系统提供严谨的安全性评估。项目页面:https://shanpoyang654.github.io/InstaDrive/page.html。
原文摘要 · Abstract (English)
Autonomous driving relies on robust models trained on high-quality, large-scale multi-view driving videos. While world models offer a cost-effective solution for generating realistic driving videos, they struggle to maintain instance-level temporal consistency and spatial geometric fidelity. To address these challenges, we propose InstaDrive, a novel framework that enhances driving video realism through two key advancements: (1) Instance Flow Guider, which extracts and propagates instance features across frames to enforce temporal consistency, preserving instance identity over time. (2) Spatial Geometric Aligner, which improves spatial reasoning, ensures precise instance positioning, and explicitly models occlusion hierarchies. By incorporating these instance-aware mechanisms, InstaDrive achieves state-of-the-art video generation quality and enhances downstream autonomous driving tasks on the nuScenes dataset. Additionally, we utilize CARLA's autopilot to procedurally and stochastically simulate rare but safety-critical driving scenarios across diverse maps and regions, enabling rigorous safety evaluation for autonomous systems. Our project page is https://shanpoyang654.github.io/InstaDrive/page.html.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。