arXiv:2503.11342cs.CV2025-03被引 1

用视觉语言模型提前识别路怒诱因并对话安抚,提升行车安全。

Road Rage Reasoning with Vision-language Models (VLMs): Task Definition and Evaluation Dataset

  • 构建路怒推理任务与标注数据集,评估模型对驾驶场景的理解能力。
  • 现有视觉语言模型在场景理解与物体空间关系识别上表现不佳。
  • 适合研究智能交通、人机交互与情绪感知的学者参考。

路怒由交通拥堵、危险驾驶等驾驶相关刺激引发,严重威胁道路安全。以往研究多聚焦于情绪抑制,缺乏主动预防能力。随着视觉语言模型(VLMs)的发展,有望通过视觉识别触发事件,并在驾驶员情绪升级前进行对话安抚。为此,我们提出路怒推理任务,构建精细标注的测试数据集与评估指标,用于评估主流VLMs在场景理解、事件识别及路怒推理方面的能力。结果表明,当前VLMs在视觉模态的场景理解以及文本模态中物体空间关系的把握上存在显著不足。提升这些能力将极大促进以源头干预为核心的路怒调控应用。

原文摘要 · Abstract (English)

Road rage, triggered by driving-related stimuli such as traffic congestion and aggressive driving, poses a significant threat to road safety. Previous research on road rage regulation has primarily focused on response suppression, lacking proactive prevention capabilities. With the advent of Vision-Language Models (VLMs), it has become possible to reason about trigger events visually and then engage in dialog-based comforting before drivers' anger escalates. To this end, we propose the road rage reasoning task, along with a finely annotated test dataset and evaluation metrics, to assess the capabilities of current mainstream VLMs in scene understanding, event recognition, and road rage reasoning. The results indicate that current VLMs exhibit significant shortcomings in scene understanding within the visual modality, as well as in comprehending the spatial relationships between objects in the textual modality. Improving VLMs' performance in these areas will greatly benefit downstream tasks like antecedent-focused road rage regulation.

路怒识别视觉语言模型情绪感知智能交通

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。