用视觉语言模型实时识别危险驾驶并生成情绪化反馈,提升行车安全与体验。
Vision-Language Assistant for Emotional Reactions to Risky Driving
- 通过YOLOv8检测突发变道等高危行为,提取速度、距离等关键指标。
- 结合用户偏好生成中性、幽默或分析型语音反馈,最高评分达4.29/5。
- 适合智能座舱、自动驾驶人机交互,尤其关注驾驶员情绪体验。
本研究提出一种视觉-语言框架Keep Yelling Assistant(KYA),用于实时检测危险驾驶行为并生成情感化响应,以增强驾驶员意识与舒适感。现有视觉-语言模型虽具备感知与推理能力,但缺乏对情绪维度和真实用户体验的考量。KYA利用YOLOv8变体检测附近车辆及突发变道等高危动作,提取相对距离、速度与预估到达时间,并将其标准化为结构化行为日志。语言模块基于用户设定的情感风格(如中性、幽默、分析)输入至ChatGPT-4o、Claude 3、Gemini 2.5和Copilot等大语言模型,生成对应口语化回应。在包含危险驾驶行为的行车记录仪视频上进行评估,并通过108名参与者的用户研究验证。结果显示所有模型均获良好评价,偏好因人格设定而异;其中YOLOv8s与ChatGPT-4o组合取得最高评分4.29/5.00。该系统融合真实感知与情绪自适应对话,为智能座舱中的情感化人机交互提供新范式,有望提升传统与自动驾驶车辆的安全性、信任度与情绪福祉。
原文摘要 · Abstract (English)
This study introduces a vision-language pipeline that detects risky driving behaviors and generates emotionally expressive responses to support driver awareness and comfort. Although vision-language models have advanced perception and reasoning in autonomous driving, existing systems rarely consider the emotional dimension or real-world user experience. Keep Yelling Assistant (KYA) detects high-risk driving maneuvers in real time, such as sudden cut-ins. It then produces emotional responses through a large language model tailored to driver preferences. The framework comprises two core modules. The vision module uses YOLOv8 variants to detect nearby vehicles and identify risky behaviors such as sudden cut-ins. Key driving metrics, including relative distance, speed, and projected reach time, are extracted and normalized to produce a structured behavior log. The language module processes this log with user-defined emotional tone settings, such as neutral, humorous, and analytical, and generates verbal reactions using state-of-the-art large language models, including ChatGPT-4o, Claude 3, Gemini 2.5, and Copilot. We evaluated the proposed system using dashcam videos containing risky driving behaviors and a user study involving 108 participants. Participants selected preferred response styles, and the large language models were evaluated based on emotional alignment. All models received favorable ratings, although preferences varied across personas. Notably, the combination of YOLOv8s and ChatGPT-4o achieved the highest score of 4.29 out of 5.00. By integrating real-world perception with emotionally adaptive dialogue, KYA introduces a new paradigm for emotionally intelligent in-vehicle artificial intelligence. It offers promising directions for improving safety, trust, and emotional well-being in both conventional and autonomous vehicles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。