构建科学逆向设计基准,测试大模型在复杂场景下的设计能力。
SciDesignBench: Benchmarking and Improving Language Models for Scientific Inverse Design
- 建立520个跨14个领域的模拟器驱动任务,覆盖多种设计场景。
- 零样本模型成功率仅29.0%,长期反馈优化后性能显著提升。
- 提出新训练方法RLSF,8B模型单次设计成功率提升8-17个百分点。
科学与工程中许多关键问题属于逆向问题:给定目标结果,寻找实现它的设计方案。评估候选方案是否达标通常较容易——如计算结合能、模拟反应产率或预测药代动力学特征。但搜索组合设计空间以找到满足目标的输入则困难得多。我们提出了SciDesignBench,一个包含520个基于模拟器的任务基准,覆盖14个科学领域和五种设置:单次设计、短周期反馈、长周期优化以及种子设计优化。在共享核心的10个领域上,最佳零样本模型成功率仅为29.0%,尽管解析率较高。模拟器反馈有助于提升性能,但排行榜随迭代次数变化:Sonnet 4.5在单轮生成中最强,而Opus 4.6在20轮模拟反馈优化后表现最优。提供初始种子设计会再次改变排名,表明受限修改与无约束生成需要完全不同能力。我们随后引入RLSF(模拟器反馈训练策略),经其微调的8B模型在三个领域中单轮成功率提升8至17个百分点。这些结果表明,基于模拟器的逆向设计既是科学推理的基准,也可作为将昂贵测试期计算转化为模型权重的实用路径。
原文摘要 · Abstract (English)
Many of the most important problems in science and engineering are inverse problems: given a desired outcome, find a design that achieves it. Evaluating whether a candidate meets the spec is often routine; a binding energy can be computed, a reactor yield simulated, a pharmacokinetic profile predicted. But searching a combinatorial design space for inputs that satisfy those targets is fundamentally harder. We introduce SciDesignBench, a benchmark of 520 simulator-grounded tasks across 14 scientific domains and five settings spanning single-shot design, short-horizon feedback, long-horizon refinement, and seed-design optimization. On the 10-domain shared-core subset, the best zero-shot model reaches only 29.0% success despite substantially higher parse rates. Simulator feedback helps, but the leaderboard changes with horizon: Sonnet 4.5 is strongest in one-turn de novo design, whereas Opus 4.6 is strongest after 20 turns of simulator-grounded refinement. Providing a starting seed design reshuffles the leaderboard again, demonstrating that constrained modification requires a fundamentally different capability from unconstrained de novo generation. We then introduce RLSF, a simulator-feedback training recipe. An RLSF-tuned 8B model raises single-turn success rates by 8-17 percentage points across three domains. Together, these results position simulator-grounded inverse design as both a benchmark for scientific reasoning and a practical substrate for amortizing expensive test-time compute into model weights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。