将视觉语言理解与生成式世界模型结合,实现端到端自动驾驶决策。
LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving
- 融合大模型语义理解与生成式场景预测,统一感知与规划。
- 在封闭环路评测中显著优于现有方法,尤其在罕见场景下表现更鲁棒。
- 支持低延迟在线规划与离线视频生成,适合复杂交通场景研究。
近年来自动驾驶技术取得显著进展,但对长尾及开放世界场景的泛化能力仍是大规模部署的主要瓶颈。部分工作利用大语言模型(LLM)和视觉语言模型(VLM)进行视觉-语言理解与推理,使车辆在生成动作时能解读罕见且高风险情境;另一些研究则探索生成式世界模型,捕捉驾驶场景的时空演化,使智能体能在行动前预演可能的未来。受人类智能中理解与想象统一的启发,我们提出首个融合LLM多模态理解与生成式世界模型的端到端闭环驾驶框架——LMGenDrive。给定多视角摄像头输入与自然语言指令,该框架可同时生成未来驾驶视频与控制信号。此设计带来互补优势:视频预测增强时空场景建模,而LLM提供强语义先验与指令对齐能力。我们进一步提出渐进式三阶段训练策略,从视觉预训练逐步过渡至多步长时距驾驶,提升稳定性与性能。LMGenDrive支持低延迟在线规划与自回归离线视频生成。实验表明,其在挑战性闭环基准上显著优于已有方法,在指令遵循、时空理解及罕见场景鲁棒性方面均有明显提升。结果表明,统一多模态理解与生成是构建更具泛化性与鲁棒性的具身决策系统的重要方向。
原文摘要 · Abstract (English)
Recent years have seen remarkable progress in autonomous driving, yet generalization to long-tail and open-world scenarios remains a major bottleneck for large-scale deployment. To address this challenge, some works use LLMs and VLMs for vision-language understanding and reasoning, enabling vehicles to interpret rare and safety-critical situations when generating actions. Others study generative world models to capture the spatio-temporal evolution of driving scenes, allowing agents to imagine possible futures before acting. Inspired by human intelligence, which unifies understanding and imagination, we explore a unified model for autonomous driving. We present LMGenDrive, the first framework that combines LLM-based multimodal understanding with generative world models for end-to-end closed-loop driving. Given multi-view camera inputs and natural-language instructions, LMGenDrive generates both future driving videos and control signals. This design provides complementary benefits: video prediction improves spatio-temporal scene modeling, while the LLM contributes strong semantic priors and instruction grounding from large-scale pretraining. We further propose a progressive three-stage training strategy, from vision pretraining to multi-step long-horizon driving, to improve stability and performance. LMGenDrive supports both low-latency online planning and autoregressive offline video generation. Experiments show that it significantly outperforms prior methods on challenging closed-loop benchmarks, with clear gains in instruction following, spatio-temporal understanding, and robustness to rare scenarios. These results suggest that unifying multimodal understanding and generation is a promising direction for more generalizable and robust embodied decision-making systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。