arXiv:2601.08434cs.ROcs.AI2026-01被引 1

用双驱动框架融合语义理解与实时决策,提升自动驾驶持续学习能力

Large Multimodal Models for Embodied Intelligent Driving: The Next Frontier in Self-Driving?

  • 结合大模型语义理解与强化学习策略优化,实现双轮驱动决策
  • 实验验证在变道规划任务中优于单一模型方法
  • 适合关注自动驾驶持续学习与端到端智能的科研人员

大型多模态模型(LMMs)为克服传统模块化自动驾驶在开放世界场景中环境理解与逻辑推理能力不足的问题提供了新可能。同时,具身人工智能通过闭环交互支持策略优化,推动自动驾驶向具身智能(EI)驾驶演进。然而,仅依赖LMMs进行决策会限制其联合规划能力。本文提出一种语义与策略双驱动的混合决策框架,融合LMMs进行语义理解与认知表征,以及深度强化学习(DRL)实现实时策略优化。文章首先阐述具身智能驾驶与LMMs的基础原理,探讨该框架带来的新机遇,包括潜在优势与典型应用场景。通过案例研究验证了该框架在完成变道规划任务中的性能优越性。最后,识别出若干未来研究方向,以进一步推动具身智能驾驶发展。

原文摘要 · Abstract (English)

The advent of Large Multimodal Models (LMMs) offers a promising technology to tackle the limitations of modular design in autonomous driving, which often falters in open-world scenarios requiring sustained environmental understanding and logical reasoning. Besides, embodied artificial intelligence facilitates policy optimization through closed-loop interactions to achieve the continuous learning capability, thereby advancing autonomous driving toward embodied intelligent (El) driving. However, such capability will be constrained by relying solely on LMMs to enhance EI driving without joint decision-making. This article introduces a novel semantics and policy dual-driven hybrid decision framework to tackle this challenge, ensuring continuous learning and joint decision. The framework merges LMMs for semantic understanding and cognitive representation, and deep reinforcement learning (DRL) for real-time policy optimization. We start by introducing the foundational principles of EI driving and LMMs. Moreover, we examine the emerging opportunities this framework enables, encompassing potential benefits and representative use cases. A case study is conducted experimentally to validate the performance superiority of our framework in completing lane-change planning task. Finally, several future research directions to empower EI driving are identified to guide subsequent work.

自动驾驶多模态模型强化学习具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。