评测科学分析模型需看数据到结论全程支持,而非只看最终结果。
SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores

- 构建多模态证据链,要求代码、结果、图表与结论全程可执行且互证。
- 11个模型在完整评估中仅62.6%通过软证据链,18.0%达成严格全链成功。
- 发现结论一致≠科学正确,必须检查每一步数据到结论的支撑关系。
科学编码代理生成相互依赖的代码、结果、图表和论断,但仅评估最终输出无法判断其结论是否具备科学依据。本文提出基于证据的多模态科学分析范式,要求代理在同一运行中生成可执行分析,并以结果与可视化支持论断。我们引入SciRIGOR评估框架与基准,包含来自六个领域、17个子领域的100个真实科学案例。该框架重构带类型证据图,区分成果保真度与关系有效性,评分时追踪完整的论断-支持路径,并定位首个不成立的关系。源码驱动的替代路径允许科学等价的分析与可视化。评估11种代理/模型配置发现:在全基准测试中,论断与忠实及非忠实结果的一致率接近(91.8%对比91.0%),但无一系统超过62.6%的软证据链得分或18.0%的严格全链成功率。结果表明,内部一致性不能保证科学正确性,评估必须验证从数据到论断的完整路径。
原文摘要 · Abstract (English)
Scientific coding agents produce interdependent code, results, figures, and claims, yet evaluating final outputs alone does not establish whether their conclusions are scientifically supported. We formulate evidence-grounded multimodal scientific analysis, requiring agents to produce executable analyses and claims supported by results and visualizations from the same run. We introduce SciRIGOR, an evaluation framework and benchmark comprising 100 cases from scientific articles across six domains and 17 subfields. The framework reconstructs typed evidence graphs, separates artifact fidelity from relational validity, and scores complete claim-support paths while localizing the earliest unsupported relation. Source-grounded alternative paths accommodate scientifically equivalent analyses and visualizations. We evaluate 11 agent/model configurations. On full-benchmark runs, claims agree with faithful and unfaithful results at nearly identical rates (91.8% versus 91.0%). Yet no system exceeds 62.6% on the soft evidence-chain score or 18.0% strict whole-chain success. These findings show that internal coherence does not establish scientific correctness: evaluation must verify support along the complete data-to-claim path.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。