用多模态大模型提升自动驾驶感知安全,让机器像人一样预判风险。
DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models
- 用定制数据集微调多模态大模型,增强对驾驶场景的理解能力。
- 闭合式问答准确率提升11.8%,开放式问答得分提高12.0%,推理仅需0.59秒。
- 已在中加两地实测,识别出人类司机也难察觉的安全隐患。
人类驾驶员具备空间与因果智能,能感知驾驶场景、预判危险并应对动态环境。相比之下,自动驾驶车辆缺乏此类能力,难以应对感知相关的预期功能安全性(SOTIF)风险,尤其在复杂或不可预测的路况下。为填补这一差距,我们提出在专为捕捉感知相关SOTIF场景设计的定制数据集上,微调多模态大语言模型(MLLMs)。基准测试显示,微调后的MLLM在闭合式视觉问答(VQA)准确率上提升11.8%,开放式VQA得分提高12.0%,同时保持实时性能,每张图像平均推理时间为0.59秒。通过加拿大与中国的真实案例研究验证了该方法的有效性,微调模型成功识别出甚至经验丰富的驾驶员也难以察觉的安全风险。本工作是首次将领域特定的MLLM微调应用于自动驾驶SOTIF领域的尝试。相关数据集与资源已公开于github.com/s95huang/DriveSOTIF.git。
原文摘要 · Abstract (English)
Human drivers possess spatial and causal intelligence, enabling them to perceive driving scenarios, anticipate hazards, and react to dynamic environments. In contrast, autonomous vehicles lack these abilities, making it challenging to manage perception-related Safety of the Intended Functionality (SOTIF) risks, especially under complex or unpredictable driving conditions. To address this gap, we propose fine-tuning multimodal large language models (MLLMs) on a customized dataset specifically designed to capture perception-related SOTIF scenarios. Benchmarking results show that fine-tuned MLLMs achieve an 11.8\% improvement in close-ended VQA accuracy and a 12.0\% increase in open-ended VQA scores compared to baseline models, while maintaining real-time performance with a 0.59-second average inference time per image. We validate our approach through real-world case studies in Canada and China, where fine-tuned models correctly identify safety risks that challenge even experienced human drivers. This work represents the first application of domain-specific MLLM fine-tuning for SOTIF domain in autonomous driving. The dataset and related resources are available at github.com/s95huang/DriveSOTIF.git
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。