用合成高风险数据提升视觉语言模型的驾驶风险预测能力
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
- 基于鸟瞰图模拟三类风险,生成可直接用于训练的高风险数据
- 在多个基线模型上将事故识别准确率从27.13%提升至88.03%
- 零样本测试下真实场景表现从29.42%升至68.50%,适合自动驾驶安全研究
自动驾驶虽取得显著进展,但在长尾场景中,由于动态环境不确定性及数据覆盖不足,准确预测自车未来运动的安全性仍面临挑战。本文探索通过合成高风险运动数据增强视觉语言模型(VLM)的运动风险预测能力。提出一种基于鸟瞰图(BEV)的运动仿真方法,从自车、其他车辆和环境三个维度建模风险,生成可即插即用的高风险数据集DriveMRP-10K。设计了不依赖特定VLM的运动风险估计框架DriveMRP-Agent,引入全局上下文、自车视角与轨迹投影的信息注入策略,使VLM能有效推理运动路径点与环境的空间关系。大量实验表明,使用DriveMRP-10K微调后,多个VLM基线的运动风险预测性能显著提升,事故识别准确率从27.13%跃升至88.03%。在自研的真实高风险运动数据集上进行零样本评估,准确率从基线模型的29.42%提升至68.50%,展现出强泛化能力。
原文摘要 · Abstract (English)
Autonomous driving has seen significant progress, driven by extensive real-world data. However, in long-tail scenarios, accurately predicting the safety of the ego vehicle's future motion remains a major challenge due to uncertainties in dynamic environments and limitations in data coverage. In this work, we aim to explore whether it is possible to enhance the motion risk prediction capabilities of Vision-Language Models (VLM) by synthesizing high-risk motion data. Specifically, we introduce a Bird's-Eye View (BEV) based motion simulation method to model risks from three aspects: the ego-vehicle, other vehicles, and the environment. This allows us to synthesize plug-and-play, high-risk motion data suitable for VLM training, which we call DriveMRP-10K. Furthermore, we design a VLM-agnostic motion risk estimation framework, named DriveMRP-Agent. This framework incorporates a novel information injection strategy for global context, ego-vehicle perspective, and trajectory projection, enabling VLMs to effectively reason about the spatial relationships between motion waypoints and the environment. Extensive experiments demonstrate that by fine-tuning with DriveMRP-10K, our DriveMRP-Agent framework can significantly improve the motion risk prediction performance of multiple VLM baselines, with the accident recognition accuracy soaring from 27.13% to 88.03%. Moreover, when tested via zero-shot evaluation on an in-house real-world high-risk motion dataset, DriveMRP-Agent achieves a significant performance leap, boosting the accuracy from base_model's 29.42% to 68.50%, which showcases the strong generalization capabilities of our method in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。