构建首个面向科学发现的自动化评估基准,测试AI能否从数据中自主提出可信结论。
TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

- 设计40个盲测任务,仅提供目标与数据,隐藏结论和分析路径,逼迫AI自主发现
- 基于6维29项证据成熟度指标,用LLM自动评分,支持可复现评估
- 发现现有编码代理虽能写报告,但缺乏控制、鲁棒性等关键科学判断能力
自主编程代理正被视作人工智能科学家系统,用于分析数据并撰写研究报告,但执行预定分析不同于真正发现。现有基准围绕隐藏目标研究设计,以结果复现为奖励。我们提出TruthInsightBench,一个面向发现的基准:包含40个来自10个科学领域的同行评审研究的盲测任务,仅暴露中性科学目标和冻结数据;源结论、预期值与分析路径均保密,要求代理自行判断数据支持的主张。固定基于LLM的评判器在六个维度上对代理提出的主张进行证据成熟度评分,共29项基于证据的评估条目,实现自动化、确定性聚合,无需人工评分,支持随代理演进而重复评估。在单一基线模型上,四个编码代理得分稳定于58.4至60.3(满分100),无统计显著差异:它们能胜任分析执行与文档撰写,具备较强证据可审计性和新颖性,但普遍缺乏建立可信主张所必需的判别性行为(如控制、鲁棒性、可证伪性、跨数据集泛化)。瓶颈在于科学判断力而非编码能力,真正的发现仍遥不可及。TruthInsightBench将此差距变为可测量的目标;数据与评分代码见https://github.com/TruthInsight-stack/TruthInsightBench。
原文摘要 · Abstract (English)
Autonomous coding agents are increasingly proposed as AI-scientist systems that conduct analyses and write research reports, but executing a prescribed analysis is not the same as making a discovery. Existing benchmarks are configured for reproduction: tasks, data, and rubrics are built around a hidden target study, and recovery of its result is rewarded. We present TruthInsightBench, a benchmark configured for discovery. Its 40 blind tasks, drawn from 40 peer-reviewed studies across 10 scientific domains, expose only a neutral scientific objective and frozen data; source conclusions, expected values, and analysis paths are withheld, leaving the agent to determine what claim the data support. A fixed LLM-based judge scores the evidentiary maturity of an agent's own claims along six dimensions, operationalized as 29 artifact-grounded items, with automated, deterministic aggregation and no per-instance human grading, so evaluation can be repeated automatically as agents evolve. On one frozen base model, four coding agents form a narrow plateau (58.4-60.3 of 100) with no statistically reliable pairwise separation: they execute and document analyses competently, with comparatively strong evidence auditability and novelty, but largely lack the discriminating acts that establish a trustworthy claim (controls, robustness, falsifiability, and cross-dataset generalization). The bottleneck is scientific judgment rather than coding, and genuine discovery remains out of reach. TruthInsightBench makes this gap a measurable target; data and scoring code are at https://github.com/TruthInsight-stack/TruthInsightBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。