用可执行评分程序动态优化实验选择,提升反应优化效率。
CARE: Context-Aware Ranking Evolution with Executable Scoring Programs for Budgeted Reaction Optimization
- 让大模型生成可执行的评分程序,自动排序待测反应条件。
- 在8个任务上实现最低归一化遗憾和最高15次内成功概率。
- 适合需要高效实验设计的化学合成与自动化研发场景。
高通量实验虽能评估大量反应条件,但组合空间仍超出可用实验预算,导致实验选择成为序列决策问题:每次需在结果未知时从有限观测中选择新条件。大语言模型(LLM)可表达任务特定的选择逻辑,但直接推荐既非持久可验证的可执行对象,也非独立可审计的决策规则。我们提出CARE,一种基于参考条件的控制器,将程序生成与实验选择分离。一个LLM编写可执行评分程序对剩余条件进行排名,而非LLM的参考策略提供数值候选与支持摘要。CARE从程序生成可选方案,通过参考条件驱动的干预门控将其与参考方案比较,并在结果揭示前记录决策。每次新结果更新控制器状态,可触发程序保留、修订或重生成。这种由结果引导的程序演化在不更新LLM参数的情况下改变评分逻辑。在8个反应优化任务上,使用30个种子进行离线回放测试,CARE在归一化遗憾、最佳进度归一化AUC及15次内前1%成功率方面均优于其他方法,表明基于LLM生成的评分程序作为参考条件优化器的组件,比作为独立选择器更有效。
原文摘要 · Abstract (English)
High-throughput experimentation can evaluate many reaction conditions, yet combinatorial condition spaces still exceed the available experiment budget. This makes experiment selection a sequential decision problem: each new condition must be chosen from limited observations before its outcome is known. LLMs can express task-specific selection logic. A direct recommendation, however, is neither a persistent executable object that can be validated and revised nor an independently auditable decision rule. We introduce CARE, a reference-conditioned controller that separates program synthesis from experiment selection. An LLM writes an executable scoring program that ranks the remaining conditions, while a non-LLM reference policy supplies a numerical candidate and support summary. CARE forms an optional alternative from the program, applies a reference-conditioned intervention gate to compare it with the reference, and records the decision before the selected outcome is revealed. Each new outcome updates controller state and can trigger retention, revision, or regeneration of the active program. This outcome-guided program evolution changes the scoring logic without updating the LLM parameters. In matched offline replay with 30 seeds on eight reaction-optimization tasks, CARE attains the lowest normalized regret, the highest normalized best-so-far AUC, and the highest Top-1% Success@15 among the evaluated methods. These results support using a scoring program written by an LLM as one component of a reference-conditioned optimizer rather than as a standalone experiment selector.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。