arXiv:2504.13068cs.CLcs.AI2025-04被引 4

模型准确率高未必符合专家判断,大模型更懂事故描述逻辑。

Accuracy is Not Agreement: Expert-Aligned Evaluation of Crash Narrative Classification Models

  • 用专家标注对比五种深度学习模型和四个大模型的表现
  • 准确率高的模型反而与专家意见分歧更大,大模型虽准确率低但更贴近专家
  • 发现专家判断依赖上下文和时间线索,非单一关键词匹配

本研究探讨深度学习模型在交通事故叙述分类中的准确率与专家一致性之间的关系。我们评估了五种深度学习模型(包括BERT变体、USE及零样本分类器)以及四种大语言模型(GPT-4、LLaMA 3、Qwen、Claude)在专家标签和事故叙述上的表现。结果表明存在反向关系:技术准确率较高的模型往往与人类专家意见一致度更低,而大语言模型尽管准确率较低,却展现出更强的专家对齐能力。通过Cohen's Kappa和主成分分析(PCA)量化并可视化模型与专家的一致性,利用SHAP分析解释误分类原因。结果显示,专家一致的模型更依赖上下文和时间线索,而非特定地点关键词。研究指出,仅依赖准确率不足以评估安全关键型NLP任务,应将专家一致性纳入评估框架,并强调大语言模型在事故分析流程中作为可解释工具的潜力。

原文摘要 · Abstract (English)

This study investigates the relationship between deep learning (DL) model accuracy and expert agreement in classifying crash narratives. We evaluate five DL models -- including BERT variants, USE, and a zero-shot classifier -- against expert labels and narratives, and extend the analysis to four large language models (LLMs): GPT-4, LLaMA 3, Qwen, and Claude. Our findings reveal an inverse relationship: models with higher technical accuracy often show lower agreement with human experts, while LLMs demonstrate stronger expert alignment despite lower accuracy. We use Cohen's Kappa and Principal Component Analysis (PCA) to quantify and visualize model-expert agreement, and employ SHAP analysis to explain misclassifications. Results show that expert-aligned models rely more on contextual and temporal cues than location-specific keywords. These findings suggest that accuracy alone is insufficient for safety-critical NLP tasks. We argue for incorporating expert agreement into model evaluation frameworks and highlight the potential of LLMs as interpretable tools in crash analysis pipelines.

自然语言处理专家对齐大模型应用事故分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。