arXiv:2509.25868cs.CL2025-09Conference of the …被引 3

评测大模型科学幻觉,发现其错误预测多与真实错误无关。

ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations

  • 构建1001组专家标注的问答对,定位错误片段并标记位置。
  • 61%的错误预测与真实错误语义无关,且大小模型均存在此问题。
  • 对比判断比独立检测更难,挑战大模型自评可靠性。

大型语言模型(LLM)在科学幻觉中的机制仍不明确。我们提出ReFACT(Reddit虚假与正确文本),一个包含1,001组由专家标注的问答对的基准,其错误片段标注源自Reddit的r/AskScience板块。评估9个最先进的LLM发现两个关键局限:第一,模型表现出主导性的“显著干扰项”失败模式——61%的错误片段预测与真实错误在语义上无关。这一现象在所有模型规模(1B至70B)下均持续存在,表明单纯扩大规模无法解决根本性的语义锚定缺陷。第二,我们发现对比判断比独立检测更困难,即使GPT-4o的F1分数也从0.67降至0.53。这些结果直接挑战了以大模型作为评判者来评估科学事实可靠性的可行性。代码与数据已公开于https://github.com/ddz5431/ReFACT。

原文摘要 · Abstract (English)

The mechanisms underlying scientific confabulation in Large Language Models (LLMs) remain poorly understood. We introduce ReFACT (Reddit False And Correct Texts), a benchmark of 1,001 expert-annotated question-answer pairs with span-level error annotations derived from Reddit's r/AskScience. Evaluating 9 state-of-the-art LLMs reveals two critical limitations. First, models exhibit a dominant "salient distractor" failure mode: 61% of incorrect span predictions are semantically unrelated to actual errors. Crucially, this pattern persists across all model scales (1B to 70B), indicating a fundamental semantic grounding deficit that scaling alone fails to resolve. Second, we find that comparative judgment is paradoxically harder than independent detection, even GPT-4o's F1 score drops from 0.67 to 0.53 when comparing answers side-by-side. These findings directly challenge the reliability of LLM-as-Judge paradigms for scientific factuality. Code and data are released at https://github.com/ddz5431/ReFACT.

幻觉检测大模型评估科学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。