arXiv:2608.23047cs.CLcs.AI2026-08

对比人类与大模型在科学事实核查中的推理路径,发现各有优劣。

Beyond Verdicts: A Graph-Based Analysis of Human and LLM Reasoning in Scientific Fact-Checking

论文配图:Beyond Verdicts: A Graph-Based Analysis of Human and LLM Reasoning in Scientific Fact-Checking
图 1 · 摘自论文原文
  • 构建基于图的推理框架,对齐人类与大模型的错误推理细节
  • 三款模型中,Qwen3-32B犯错最少,GPT-5最接近人类思路
  • 即使推理路径不同,部分模型仍能给出有效解释,适合可信度评估

虚假信息若引用真实论文,可能因曲解研究内容而造成严重危害。现有基于大语言模型(LLM)的自动事实核查系统虽可判断模型输出‘错误’结论并生成解释,但通常无法判断其推理路径是否与人类专家一致,或是否通过另一条合理路径得出相同结论。本文提出一种基于图的分析框架(类型化推理图),用于比较人类与LLM在科学事实核查中的推理路径。基于生物医学误导信息中的谬误研究(MISSCIPLUS, Glockner et al., 2025),将每条解释建模为连接错误主张、研究背景、研究发现、支持谬误的前提及谬误标签的推理图。该表示支持人类与LLM推理在特定谬误子图层面的一一对应。对于不匹配人类路径的LLM推理,我们验证其在引用研究中的依据性、与主张的相关性以及对结论的充分性。基于84个来自MISSCIPLUS的虚假主张,我们在提示词和证据设置下评估了GPT-5、Claude Opus 4.7和Qwen3-32B。结果显示:三者表现维度各异——Qwen3-32B的结论错误率最低,GPT-5的人类推理对齐度最高,Claude Opus 4.7虽然结论预测能力较弱,但在成功案例中常表现出有效推理。

原文摘要 · Abstract (English)

Misinformation that cites legitimate papers can be especially harmful when it distorts what those studies actually report. While existing automatic fact-checking systems based on large language models (LLMs) can assess whether a model assigns an Incorrect verdict and can gen- erate explanations for that decision, they typi- cally do not indicate whether the model follows the same reasoning path as human experts or arrives at the verdict through a different but still valid path. In this work, we introduce a graph- based framework (typed reasoning graph) for comparing human and LLM reasoning paths in scientific fact-checking. Building on prior work on fallacious reasoning in biomedical misinformation, MISSCIPLUS (Glockner et al., 2025), we model each explanation as a rea- soning graph that links the false claim to the relevant study context, study findings, fallacy- supporting premises, and fallacy labels. This representation enables one-to-one alignment of human and LLM reasoning at the level of fallacy-specific sub-graphs. For non-human- aligned LLM paths, we validate grounding in the cited study, relevance to the claim, and suf- ficiency for the verdict. Using 84 false claims from MISSCIPLUS, we evaluate GPT-5, Claude Opus 4.7, and Qwen3-32B across prompt and evidence settings. Results show distinct perfor- mance dimensions: Qwen3-32B has the lowest verdict failure rate, GPT-5 the highest human alignment, and Claude Opus 4.7 weak verdict prediction but often valid reasoning in success- ful cases

事实核查推理图大模型评估科学可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。