分析人类写作中的事实错误,发现新类型并验证大模型检测能力有限。
An Empirical Analysis of Factual Errors in Human-Written Text and its Application

- 基于报纸修正数据构建人类错误分类体系,发现汉字误用、数字量词错等特有类型。
- 高阶大模型在合成测试中仅达52%的单词级F1,表明人类事实错误检测难度高。
- 为真实文本纠错与大模型可信度评估提供新基准,适合语言模型安全研究者。
事实错误检测(FED)是识别文本中事实性错误片段的重要任务。随着大语言模型(LLM)的兴起,研究焦点转向生成文本中的幻觉及其检测,而人类写作中的事实错误检测被相对忽视。为弥补这一空白,我们通过分析报纸文章的修改记录,构建了人类引发事实错误的分类体系。分析发现存在汉字误用、数字量词错误等典型类别,现有幻觉评测数据集未覆盖。基于该分类体系,我们在合成真实测试样本和真实修正数据上评估了原始大模型的FED能力。实验显示,即使高性能模型GPT-5.4在合成数据上的单词级F1也仅为52%,凸显任务挑战性。进一步按检测难度分析揭示当前FED技术的局限。
原文摘要 · Abstract (English)
Factual Error Detection (FED), which is the task of identifying factually incorrect spans in a given text, has long been recognized as an important research problem. However, with the rapid rise of large language models (LLMs), research attention has shifted toward factual errors specific to LLM-generated text (hallucinations) and their detection. As a result, the detection of factual errors in human-written text has been relatively neglected. To address this gap, we first distill a taxonomy of human-induced factual errors by analyzing corrections of newspaper articles, a representative source of text that is guaranteed to be human-written and contains few grammatical errors. Our analysis revealed that there are characteristic categories such as kanji misconversions and numeral classifier errors, which are not focused in existing hallucination benchmarks. Based on the taxonomy, we then evaluate the FED capability of vanilla LLMs on synthesized realistic test cases and real corrections. Experimental results demonstrated that even high-performance LLMs such as GPT-5.4 achieved only word-level F1 score of 52% on the synthetic evaluation data, highlighting the task difficulty. Furthermore, a detailed analysis by detection difficulty revealed the current state of FED.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。