测试了AI在医学等高风险领域合成科学结论的能力,发现效果仍不理想。
Can AI Agents Synthesize Scientific Conclusions?

- 构建9110个问题的基准测试集,用原子事实评估结论正确性和完整性。
- 最佳模型在严格隔离环境下仅达0.337的事实F1,远未达标。
- 揭示现有AI工具常生成不完整或自相矛盾的结论,适合关注可信AI的研究者。
科学AI代理越来越多地用于检索证据、跨源推理并合成决策性结论,但在医疗等高风险领域的合成能力尚不明确。本文提出SciConBench,一个包含9.11K个问题和专家撰写结论的大规模实时基准,基于专家验证的自动评估流程,将结论分解为原子事实,通过事实精确率和召回率衡量正确性与全面性。为防止数据泄露,引入SciConHarness——一种受控环境下的评估框架,限制代理的网络交互以确保有效测量。评估8个前沿模型及深度研究代理后发现,事实质量普遍偏低:在清洁室设置下,表现最好的代理仅获得0.337的事实F1。清洁室设置始终使性能低于非受限评估,表明数据泄露会夸大模型真实合成能力。最后审计了面向消费者的代理(如Google AI Overview、OpenEvidence),发现其即使面对已知正确答案,仍常生成不完整或矛盾的结论。总体表明,可靠科学结论合成仍是开放挑战,且清洁室评估对公正评估开放域AI代理至关重要。
原文摘要 · Abstract (English)
Scientific AI agents increasingly retrieve evidence, reason across sources, and synthesize conclusions used in consequential decisions. Yet, their ability to do so in high-stakes domains such as health remains unclear. We introduce SciConBench, a large-scale live benchmark of 9.11K questions and expert-written conclusions from systematic reviews to evaluate open-domain scientific conclusion synthesis. The benchmark draws on an expert-validated automated evaluation pipeline that decomposes conclusions into atomic facts and measures correctness and comprehensiveness via factual precision and recall. To mitigate data leakage, we further introduce SciConHarness, a clean-room evaluation harness that equips agents with controlled web interaction to ensure valid measurement. Evaluating 8 frontier models and deep research agents, we find that factual quality remains low: under clean-room settings, the best agent achieves only a factual F1 of 0.337. Our clean-room setting consistently reduces performance relative to unconstrained evaluation, suggesting that leakage inflates estimates of models' true synthesis capabilities. Finally, we audit consumer-facing agents (e.g., Google AI Overview, OpenEvidence) and find they frequently generate incomplete and sometimes contradictory conclusions, even when the ground-truth answer is available. Overall, our results show that reliable synthesis of scientific conclusions remains an open challenge, and that clean-room evaluation is essential for assessing open-domain AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。