arXiv:2511.13909cs.CV2025-11

评测大模型对交通规则图示的理解能力,发现其远不如人类。

Mind the Gap: Evaluating LLM Understanding of Human-Taught Road Safety Principles

  • 用教材插图构建数据集,零样本测试多模态大模型
  • 模型在安全推理任务上表现差,准确率显著低于人类
  • 揭示人与模型理解差异,为未来研究提供方向

遵守交通规则对人类和自动驾驶的AI系统都至关重要。本文评估多模态大语言模型(LLMs)对道路安全概念的理解能力,特别是通过示意图和图示进行理解。我们从教材中收集了包含交通标志和安全规范的图像,构建了一个试点数据集,并在零样本设置下评估模型性能。初步结果显示,这些模型在安全推理方面表现不佳,暴露出人类学习与模型解读之间的差距。本文还进一步分析了这些性能差距,为后续研究提供依据。

原文摘要 · Abstract (English)

Following road safety norms is non-negotiable not only for humans but also for the AI systems that govern autonomous vehicles. In this work, we evaluate how well multi-modal large language models (LLMs) understand road safety concepts, specifically through schematic and illustrative representations. We curate a pilot dataset of images depicting traffic signs and road-safety norms sourced from school text books and use it to evaluate models capabilities in a zero-shot setting. Our preliminary results show that these models struggle with safety reasoning and reveal gaps between human learning and model interpretation. We further provide an analysis of these performance gaps for future research.

大模型安全推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。