arXiv:2411.11913cs.AIcs.RO2024-11被引 27

用视觉语言模型让汽车懂用户偏好,实现实时个性化驾驶。

On-Board Vision-Language Models for Personalized Autonomous Vehicle Motion Control: System Design and Real-World Validation

  • 基于检索增强生成的内存模块,持续学习用户驾驶偏好。
  • 实测将接管率降低76.9%,兼顾安全与舒适性。
  • 首个在真实车辆中部署的端到端VLM驾驶控制系统。

个性化驾驶指自动驾驶汽车能根据个体用户的偏好和驾驶风格调整行为,同时保证安全与舒适。现有方法或难以精准捕捉用户偏好,或随用户规模扩大而计算效率下降。视觉语言模型(VLM)凭借其自然语言理解与场景推理能力,为该问题提供新解。本文提出一种轻量级车载VLM框架,在保持强推理能力的同时实现低延迟个性化控制。系统引入基于检索增强生成(RAG)的记忆模块,通过用户反馈持续学习个体驾驶偏好。经全面真实道路车辆部署与实验验证,该系统在多种场景下均能提供安全、舒适且个性化的驾驶体验,接管率最高降低76.9%。据我们所知,这是首个在真实自动驾驶车辆中实现的端到端VLM运动控制系统。

原文摘要 · Abstract (English)

Personalized driving refers to an autonomous vehicle's ability to adapt its driving behavior or control strategies to match individual users' preferences and driving styles while maintaining safety and comfort standards. However, existing works either fail to capture every individual preference precisely or become computationally inefficient as the user base expands. Vision-Language Models (VLMs) offer promising solutions to this front through their natural language understanding and scene reasoning capabilities. In this work, we propose a lightweight yet effective on-board VLM framework that provides low-latency personalized driving performance while maintaining strong reasoning capabilities. Our solution incorporates a Retrieval-Augmented Generation (RAG)-based memory module that enables continuous learning of individual driving preferences through human feedback. Through comprehensive real-world vehicle deployment and experiments, our system has demonstrated the ability to provide safe, comfortable, and personalized driving experiences across various scenarios and significantly reduce takeover rates by up to 76.9%. To the best of our knowledge, this work represents the first end-to-end VLM-based motion control system in real-world autonomous vehicles.

自动驾驶视觉语言模型个性化控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。