让自动驾驶更懂个人习惯,能自动判断风险并做出个性化决策。
PADriver: Towards Personalized Autonomous Driving
- 用多模态大模型分析视频和用户提示,闭环生成驾驶行为。
- 在250小时标注数据上表现超越现有方法,支持多种驾驶模式。
- 适合研究个性化智能驾驶的开发者与工程师使用。
本文提出PADriver,一种新型闭合回路的个性化自动驾驶(PAD)框架。基于多模态大语言模型(MLLM),PADriver以流式图像帧和个性化文本提示为输入,自动执行场景理解、危险等级评估与动作决策。预测的危险等级反映潜在动作的风险,并为最终决策提供明确参考,对应预设的个性化提示。此外,我们基于Highway-Env模拟器构建了名为PAD-Highway的闭合回路基准,用于全面评估交通规则下的决策性能。该数据集包含250小时高质量标注视频,助力个性化驾驶行为分析。在所构建基准上的实验结果表明,PADriver在不同评估指标上优于现有先进方法,并支持多种驾驶模式。
原文摘要 · Abstract (English)
In this paper, we propose PADriver, a novel closed-loop framework for personalized autonomous driving (PAD). Built upon Multi-modal Large Language Model (MLLM), PADriver takes streaming frames and personalized textual prompts as inputs. It autoaggressively performs scene understanding, danger level estimation and action decision. The predicted danger level reflects the risk of the potential action and provides an explicit reference for the final action, which corresponds to the preset personalized prompt. Moreover, we construct a closed-loop benchmark named PAD-Highway based on Highway-Env simulator to comprehensively evaluate the decision performance under traffic rules. The dataset contains 250 hours videos with high-quality annotation to facilitate the development of PAD behavior analysis. Experimental results on the constructed benchmark show that PADriver outperforms state-of-the-art approaches on different evaluation metrics, and enables various driving modes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。