arXiv:2507.08459cs.CL2025-07ACL被引 5

给大模型的回答错误分类并归因,提升故障诊断效率。

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework

  • 构建6大类15小类错误归因框架,系统化分类模型答错原因。
  • 提出AttriData数据集,包含错误归因标签与评分反馈。
  • 开发首个能同时输出评分、归因和建议的通用判别模型。

随着大语言模型在各类任务中广泛应用,主流平台每日产生海量用户-模型交互。为高效分析模型表现并诊断回答错误,亟需建立自动化框架以系统性地分类与归因错误。然而现有评估方法缺乏错误归因能力。本文提出一个涵盖6个主类别与15个次级类别的综合错误归因框架,基于此构建了专用于错误归因的AttriData数据集,包含错误归因标签、对应评分与反馈信息。进一步提出MisAttributionLLM,该模型在AttriData上微调,是首个能同时生成评分、错误归因与反馈的通用判别模型。通过大量实验与分析,验证了所提方法的有效性与鲁棒性。

原文摘要 · Abstract (English)

With the widespread application of Large Language Models (LLMs) in various tasks, the mainstream LLM platforms generate massive user-model interactions daily. In order to efficiently analyze the performance of models and diagnose failures in their answers, it is essential to develop an automated framework to systematically categorize and attribute errors. However, existing evaluation models lack error attribution capability. In this work, we establish a comprehensive Misattribution Framework with 6 primary and 15 secondary categories to facilitate in-depth analysis. Based on this framework, we present AttriData, a dataset specifically designed for error attribution, encompassing misattribution, along with the corresponding scores and feedback. We also propose MisAttributionLLM, a fine-tuned model on AttriData, which is the first general-purpose judge model capable of simultaneously generating score, misattribution, and feedback. Extensive experiments and analyses are conducted to confirm the effectiveness and robustness of our proposed method.

错误归因评估框架大模型诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。