arXiv:2606.18060cs.AIcs.CL2026-06

测试大模型科研代理识别伪科学的能力,发现多数系统难以拒绝伪科学结论。

PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience

论文配图:PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience
图 1 · 摘自论文原文
  • 构建对抗性基准PseudoBench,评估代理从实验到写作全流程的抗伪科学能力。
  • 7个顶尖代理中,拒识率仅27.4%,多数生成看似可信的伪科学报告。
  • 更强模型反而用更专业的语言包装伪科学,危害更大,需警惕部署风险。

随着基于大语言模型的智能体进入自主科研领域,其抵抗伪科学的能力日益重要。否则,这些系统可能快速生成看似合理却误导性的研究,污染学术文献并削弱公众对科学的信任。本文提出PseudoBench,一个对抗性基准,用于评估智能体在端到端科研流程(从实验设计到论文撰写)中识别与抵制伪科学叙事的能力。该基准包含跨五个领域的200组伪科学主张-证据对。测试七个最先进代理后发现,当前系统几乎零拒绝率地生成与伪科学前提一致的报告,最高抵抗率仅为27.4%。更强的代理甚至能以更精炼的科学语言包装伪科学,提升其可信度。这一结果揭示了其助长伪科学的巨大风险,亟需在广泛部署前加强科学对齐。

原文摘要 · Abstract (English)

As Large Language Model based agents enter autonomous scientific research, their ability to resist pseudoscience becomes increasingly important. Otherwise, such systems may rapidly generate plausible yet misleading studies that contaminate academic literature and erode trust in science. We present PseudoBench, an adversarial benchmark for evaluating whether agentic auto-research systems can identify and resist pseudoscientific narratives. PseudoBench contains 200 curated pseudoscientific claim-evidence pairs across five domains and evaluates agents through an end-to-end research pipeline from experiments to writing. Testing seven state-of-the-art agents, we find that current systems readily produce persuasive reports that align with pseudoscientific premises with near-zero refusal rates and the highest resistance of only 27.4%. Stronger agents risk packaging pseudoscience in more sophisticated scientific language, increasing its apparent credibility. These findings reveal an alarming capacity to fuel pseudoscience, calling for scientific alignment before widespread deployment.

伪科学检测智能体安全科研自动化科学对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。