arXiv:2605.15341cs.LGcs.AI2026-05被引 1

用轨迹评估方法发现大模型在科学设计中并不如预期高效。

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design

论文配图:LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design
图 1 · 摘自论文原文
  • 引入轨迹级评估框架,关注学习过程而非仅终点结果
  • 53%任务中结论因评估方式改变,显示传统方法会误判效率
  • 领域无关提示反而更接近文献最优解,揭示模型偏差

大语言模型在自主实验室中的应用日益广泛,普遍认为其具备领域先验和迭代反馈推理能力,可在较少迭代中收敛到优质设计。然而,现有迭代科学设计基准仅在固定时间点评分结果,未衡量学习轨迹,而轨迹才真正反映效率——每节省一次迭代即降低真实成本与时间。为此,我们考察三种评估选择如何影响对模型学习效率的判断:测什么、比什么基准、以什么为参照。提出LEAPBench:学习效率在自适应过程中的评估框架,包含55个任务,采用最佳当前值面积(AUC)轨迹指标,对比经典贝叶斯优化基线,并以已发表文献为审计依据。在8个主流大模型上的应用显示,从终点结果转向轨迹评分后,53%的任务中最佳模型决策发生变化,揭示了原方法遗漏的效率增益;且大模型并未优于经典贝叶斯基线。在16个生物学任务中,当理想奖励信号与文献最优配置一致时,领域感知提示在第30轮匹配文献最优解的准确率比领域无关提示低约10个百分点;该现象在6个文献典型配置与最优配置相悖的任务上尤为明显,领域无关提示在所有6个任务中均表现更优。此外,轨迹指标可作为可训练目标,基于该指标的离线强化学习在21个保留任务中提升了14个的表现。

原文摘要 · Abstract (English)

LLMs are increasingly deployed in autonomous laboratories, under the assumption that their domain priors and reasoning over iterative feedback let them converge on good designs in fewer iterations than feedback-only baselines. Current iterative scientific design benchmarks, however, score only outcome snapshots at fixed horizons. This leaves the learning trajectory unmeasured, even though the trajectory is what captures learning efficiency, where each iteration saved is a real saving in cost and time. Motivated by this, we examine three evaluation choices that change the conclusions one draws about LLM learning efficiency in iterative scientific design: what to measure, what baseline to compare against, and what to ground against. We introduce LEAPBench, Learning Efficiency in Adaptive Processes, a 55-task framework that pairs a best-so-far area under the curve (AUC) trajectory metric with a classical Bayesian-optimization reference and an audit grounded in published literature. Applied to eight contemporary LLMs, switching from final-outcome to trajectory scoring changes the best-model decision on 53% of tasks at matched horizons, and exposes efficiency gains overlooked by outcome-based scoring. LLMs do not outperform a classical Bayesian baseline. On 16 biology tasks where the oracle's reward signal is aligned with configurations from the published-best design, domain-aware prompting leads to LLM choices that match the published-best's approximately 10 percentage points less often than domain-agnostic prompting at iteration 30. The pattern is sharpest on 6 tasks where the literature-typical and published-best configurations diverge, and domain-agnostic prompting matches the published-best more often on all 6. The trajectory metric also doubles as a tractable training target. Offline reinforcement learning with the metric as a reward improves performance on 14 of 21 held-out tasks.

大模型评估科学设计轨迹分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。