测试大模型识别伪科学论文的能力,发现多数模型无法可靠判断科学可信度。
TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

- 用42篇撤稿/造假论文构造测试集,让模型判断其科学合理性
- 95%的非空回复中模型仍接受错误前提,超71%任务失败率
- 模型表现依赖特定高关注度话题,缺乏真实科学判断能力
大型语言模型正被提议用于科学工作流中的代理角色,但在缺乏下游验证者的情况下,其能否区分可靠与不可靠科学文献尚未直接测量。现有基准仅评估已知答案的事实性,而本文关注不同故障模式。我们引入一个包含42篇被撤稿、伪造及伪科学论文的探测语料库,并设计方法评估模型对每篇论文框架的单次响应。每个探针由目标论文近似原文提取的引言段落和一个看似合理的实验设计请求组成。探针涵盖五类主张:虚构观测、伪物理机制、魔法前提、合法化桥梁和伪实验。通过两个互补指标衡量模型是否彻底拒绝错误前提(IFR-a)以及是否在承认不可靠性的同时仍继续回应(IFR-i)。深度得分——参与深度指数(EDI)——量化模型复现论文或领域特有隐含细节的程度。在30个模型、10次重复运行下,总体IFR-a为0.93±0.004,IFR-i为0.809±0.009。模型在95%的非空响应中仍与不可接受的前提互动。所有评估模型在超过71%的任务中失败,其中22/30模型在超过90%情况下失败。拒绝对应集中于少数高知名度主题和特定探针,且在结构匹配控制下消失。结果表明,模型表现反映的是基于主题的安全行为而非稳健的认知可靠性,凸显科学部署中亟需建立防护机制。
原文摘要 · Abstract (English)
Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists. Such deployment assumes the model can distinguish reliable scientific literature from unreliable literature, a capability that has not yet been directly measured. Existing benchmarks evaluate factuality on questions with known answers; the failure mode we target here is different. We introduce a probe corpus of 42 retracted, fraudulent, and pseudoscientific papers, paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing. Each probe pairs a preamble extracted near-verbatim from the target paper with a scientifically plausible study-design request. The probes span five claim types: fabricated observation, pseudophysical mechanism, magical premise, legitimization bridge, and cargo-cult experiment. Two complementary scores measure whether a model rejects the flawed premise outright (IFR-a) and whether it recognizes the unreliability while still engaging (IFR-i). A depth score, the Engagement Depth Index (EDI), quantifies reproduction of paper- or field-specific withheld details. Across 30 models and 10 repeated runs, aggregate IFR-a is 0.93 $\pm$ 0.004 and aggregate IFR-i is 0.809 $\pm$ 0.009. Models engaged with untenable premises in 95% of all non-empty responses. Every evaluated model fails more than 71% of agentic probes, and 22 of 30 models fail more than 90% of the time. Rejections are concentrated on a small number of high-notoriety topics and specific probes, and disappear under matched-structure controls. These results are consistent with topic-keyed safety behavior rather than robust epistemic competence, and indicate an urgent need for guardrail infrastructure for scientific deployment of language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。