为大模型科学发现能力设计新评测基准,聚焦六类核心科研能力。
Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

- 按科学推理类型重构评测体系,分六类能力与五领域
- 测试15个模型,仅描述性任务表现良好,深层推理普遍失败
- 提供五阶段错误分析框架,定位模型在假设选择等环节短板
现有科学数据分析评测主要关注代码执行或流程完成,忽视科学分析支持不同类型的科学主张——假设探索、统计推断、机制解释——各自具有不同的假设和有效性标准。我们提出SDABench,一个以六种能力(描述性、探索性、推断性、预测性、因果性、机制性)为核心,覆盖五个领域(生物、化学、环境、地理、物理)的评测基准。SDABench包含527个真实数据实例(SDA-Real)和6000个合成实例(SDA-Synth),每项任务均提供多选与开放问答两种格式,由自动化流程构建。评估15个代表性LLM发现,模型在描述性分析中表现良好,但在需要假设选择、潜在过程建模或机制推理的任务上显著退化。此外,基准提供五阶段错误分析框架,揭示更先进模型虽能更好识别相关范围与变量,但仍难以选择恰当分析方法、建模变量关系并得出有效结论。
原文摘要 · Abstract (English)
Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, each with different assumptions and validity criteria. We introduce SDABench, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics). SDABench comprises 527 real-data instances (SDA-Real) and 6000 synthetic instances (SDA-Synth), each in both multiple-choice and open-ended formats, constructed through an automated pipeline. Evaluating 15 representative LLMs, we find that models handle descriptive analysis well but degrade sharply on tasks requiring assumption selection, latent-process modeling, or mechanistic reasoning. SDABench further provides a five-stage error analysis framework that locates where LLMs fail: more advanced models more reliably identify the relevant scope and variables, but still struggle to select appropriate analytical procedures, model variable relationships, and draw valid conclusions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。