arXiv:2507.08367cs.CVcs.SY2025-07

用大模型分析行车画面,评估老年人驾驶风险。

Understanding Driving Risks using Large Language Models: Toward Elderly Driver Assessment

  • 用提示词设计让大模型理解交通场景,而非简单识别物体。
  • 多轮提示使路口可视性识别召回率从21.7%提升至57.0%。
  • 适合做老年司机风险评估的辅助工具,解释性强。

本研究探讨了多模态大语言模型ChatGPT-4o在静态行车记录仪图像上进行类人化交通场景解读的潜力。聚焦于三项与老年司机评估相关的判断任务:交通密度评估、路口可视性判断和停车标志识别。这些任务需要上下文推理而非简单目标检测。采用零样本、少样本和多样本提示策略,以人工标注为标准参考,评估模型性能,使用精确率、召回率和F1分数作为指标。结果表明,提示设计显著影响表现:路口可视性召回率从零样本的21.7%提升至多样本的57.0%;交通密度一致性从53.5%提升至67.6%。停车标志检测表现出高精确率(最高86.3%),但召回率较低(约76.7%),显示模型倾向保守响应。输出稳定性分析发现,人类与模型在结构模糊场景中均存在困难。然而,模型生成的解释文本与其预测一致,增强了可解释性。结果表明,通过精心设计提示,大模型有望成为场景级驾驶风险评估的辅助工具。未来研究应探索更大数据集、多样标注者及新一代模型架构在老年司机评估中的可扩展性。

原文摘要 · Abstract (English)

This study investigates the potential of a multimodal large language model (LLM), specifically ChatGPT-4o, to perform human-like interpretations of traffic scenes using static dashcam images. Herein, we focus on three judgment tasks relevant to elderly driver assessments: evaluating traffic density, assessing intersection visibility, and recognizing stop signs recognition. These tasks require contextual reasoning rather than simple object detection. Using zero-shot, few-shot, and multi-shot prompting strategies, we evaluated the performance of the model with human annotations serving as the reference standard. Evaluation metrics included precision, recall, and F1-score. Results indicate that prompt design considerably affects performance, with recall for intersection visibility increasing from 21.7% (zero-shot) to 57.0% (multi-shot). For traffic density, agreement increased from 53.5% to 67.6%. In stop-sign detection, the model demonstrated high precision (up to 86.3%) but a lower recall (approximately 76.7%), indicating a conservative response tendency. Output stability analysis revealed that humans and the model faced difficulties interpreting structurally ambiguous scenes. However, the model's explanatory texts corresponded with its predictions, enhancing interpretability. These findings suggest that, with well-designed prompts, LLMs hold promise as supportive tools for scene-level driving risk assessments. Future studies should explore scalability using larger datasets, diverse annotators, and next-generation model architectures for elderly driver assessments.

大模型驾驶风险老年驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。