用视觉语言模型提前1-3秒预测车道保持系统故障并解释原因。
Bridging Human Oversight and Black-box Driver Assistance: Vision-Language Models for Predictive Alerting in Lane Keeping Assist Systems
- 结合车载视频与CAN数据,用可解释模型引导注意力
- 预测准确率69.8%,F1分数58.6%,生成自然语言解释
- 适合关注自动驾驶安全与可解释性的研究人员
车道保持辅助系统虽广泛应用,但因黑箱特性常发生不可预测的失效,限制了驾驶员的预判与信任。为弥合自动化与人机监督之间的鸿沟,我们提出LKAlert——一种基于视觉语言模型(VLM)的预测性预警系统,可提前1-3秒预测车道保持辅助(LKA)潜在风险。该系统融合倒车摄像头视频与车辆总线(CAN)数据,利用并行可解释模型生成的替代车道分割特征作为自动引导注意力机制。不同于传统二分类器,LKAlert不仅发出预警,还提供简洁自然语言解释,提升驾驶员情境感知与信任度。为此,我们构建了首个用于预测性与可解释性LKA故障预警的基准数据集OpenLKA-Alert,包含同步多模态输入及人工标注的时间窗理由。我们进一步提出一种通用方法框架,将代理特征引导与LoRA结合,使VLM能在不修改视觉主干的前提下推理结构化视觉上下文,适用于其他复杂黑箱系统的可解释监督。实证结果表明,系统对即将发生的LKA故障预测准确率达69.8%,F1-score为58.6%;生成的文本解释质量高(ROUGE-L 71.7),运行频率约2 Hz,具备实时车载应用潜力。研究证明了LKAlert在提升当前ADAS安全性与可用性方面的可行性,并提供了一种可扩展的视觉语言模型人机协同监督范式。
原文摘要 · Abstract (English)
Lane Keeping Assist systems, while increasingly prevalent, often suffer from unpredictable real-world failures, largely due to their opaque, black-box nature, which limits driver anticipation and trust. To bridge the gap between automated assistance and effective human oversight, we present LKAlert, a novel supervisory alert system that leverages VLM to forecast potential LKA risk 1-3 seconds in advance. LKAlert processes dash-cam video and CAN data, integrating surrogate lane segmentation features from a parallel interpretable model as automated guiding attention. Unlike traditional binary classifiers, LKAlert issues both predictive alert and concise natural language explanation, enhancing driver situational awareness and trust. To support the development and evaluation of such systems, we introduce OpenLKA-Alert, the first benchmark dataset designed for predictive and explainable LKA failure warnings. It contains synchronized multimodal inputs and human-authored justifications across annotated temporal windows. We further contribute a generalizable methodological framework for VLM-based black-box behavior prediction, combining surrogate feature guidance with LoRA. This framework enables VLM to reason over structured visual context without altering its vision backbone, making it broadly applicable to other complex, opaque systems requiring interpretable oversight. Empirical results correctly predicts upcoming LKA failures with 69.8% accuracy and a 58.6\% F1-score. The system also generates high-quality textual explanations for drivers (71.7 ROUGE-L) and operates efficiently at approximately 2 Hz, confirming its suitability for real-time, in-vehicle use. Our findings establish LKAlert as a practical solution for enhancing the safety and usability of current ADAS and offer a scalable paradigm for applying VLMs to human-centered supervision of black-box automation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。