arXiv:2506.00819cs.ROcs.AI2025-06被引 2

用双视觉语言模型让自动驾驶更懂路况,自动调整策略并保障安全。

DriveMind: A Dual Visual Language Model-based Reinforcement Learning Framework for Autonomous Driving

  • 用对比视觉语言模型逐步锚定语义,实时理解驾驶场景。
  • 动态生成提示词,应对路况变化,成功率提升超4%。
  • 可零样本迁移到真实摄像头数据,适合部署于实际道路。

端到端自动驾驶系统直接将传感器数据映射为控制指令,但存在不透明、缺乏可解释性且无形式化安全保证的问题。尽管近期基于视觉-语言的强化学习方法引入了语义反馈,却常依赖静态提示和固定目标,难以适应动态驾驶环境。本文提出DriveMind,一个统一的语义奖励框架,包含:(i) 对比视觉语言模型编码器,用于分步语义定位;(ii) 基于链式思维蒸馏微调的新颖性触发编码器-解码器,实现语义漂移时的动态提示生成;(iii) 分层安全模块,强制执行运动学约束(如速度、车道居中、稳定性);(iv) 紧凑的预测世界模型,以奖励与预期理想状态的一致性。DriveMind在CARLA Town 2中实现平均速度19.4 ± 2.3 km/h,路线完成率0.98 ± 0.03,近零碰撞,成功率优于基线超4%。其语义奖励在未见的真实行车数据上零样本泛化,分布偏移极小,展现出强大的跨域对齐能力,具备实际部署潜力。

原文摘要 · Abstract (English)

End-to-end autonomous driving systems map sensor data directly to control commands, but remain opaque, lack interpretability, and offer no formal safety guarantees. While recent vision-language-guided reinforcement learning (RL) methods introduce semantic feedback, they often rely on static prompts and fixed objectives, limiting adaptability to dynamic driving scenes. We present DriveMind, a unified semantic reward framework that integrates: (i) a contrastive Vision-Language Model (VLM) encoder for stepwise semantic anchoring; (ii) a novelty-triggered VLM encoder-decoder, fine-tuned via chain-of-thought (CoT) distillation, for dynamic prompt generation upon semantic drift; (iii) a hierarchical safety module enforcing kinematic constraints (e.g., speed, lane centering, stability); and (iv) a compact predictive world model to reward alignment with anticipated ideal states. DriveMind achieves 19.4 +/- 2.3 km/h average speed, 0.98 +/- 0.03 route completion, and near-zero collisions in CARLA Town 2, outperforming baselines by over 4% in success rate. Its semantic reward generalizes zero-shot to real dash-cam data with minimal distributional shift, demonstrating robust cross-domain alignment and potential for real-world deployment.

自动驾驶视觉语言模型强化学习安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。