用闭环自动化研究发现可泛化的分子性质预测改进
Closed-loop Auto Research for Molecular Property Prediction: Discovering and Certifying Generalizable Improvements
- 语言模型代理动态优化特征、模型和外部数据三方面
- 跨36个任务验证,部分改进在测试集上仍有效提升0.042
- 强调发现与验证分离,适合闭环机器学习系统设计
闭环自动化研究将自动化机器学习从固定数据集拟合拓展到动态调整研究流程,通过语言模型代理修改表示和模型代码,并获取外部证据。分子性质预测涉及众多小规模任务。我们探究该动作空间能否带来超越验证信号的泛化改进。在三个基准套件的36个任务上,对每个配置仅进行一次保留测试集评估,且搜索过程未读取测试标签。基于各任务最优验证轴的路由管道,在不同套件中分别获得0.013、0.011和0.042的正向测试提升,可转移轴因套件而异:TDC以数据为主,Polaris以模型为主,MoleculeNet则依赖特征与模型。最大模型搜索增益从验证集的0.041降至测试集0.003,而精选外部数据在测试集上表现为负0.019,呈现非转移特征。经污染过滤器(拒绝与测试结构重叠64%至89%的同源文件)筛选的外部数据,使CYP2C9底物性能提升0.17,半衰期提升0.08,表明此条件必要但不充分。对照实验显示,匹配试验的自动化机器学习无法复现代理的代码级干预,仅得0.006,而该管道在共享训练集上仍优于8400万参数预训练3D模型。实验限于分子性质预测,但发现与测试认证分离的方法具有领域无关意义,适用于任何以代理目标优化保留量的闭环系统。
原文摘要 · Abstract (English)
Closed-loop Auto Research extends automated machine learning from fixed-dataset fitting to changing the research workflow, with language-model agents editing representations and model code and acquiring external evidence. Molecular property prediction spans many small endpoints. We ask whether this action space yields improvements generalizing beyond the validation signal selecting them. We isolate three Auto Research axes, features, models, and external evidence, under a file-level ablation lock attributing each gain to one axis over a strong baseline. Across 36 endpoints in three benchmark suites we score each selected configuration once on a held-out test whose labels the search never read. A routed pipeline taking each endpoint's best validation axis reaches positive held-out gains of 0.013, 0.011, and 0.042, the transferable axis differing by suite, data on TDC, model on Polaris, feature and model on MoleculeNet. The largest model-search gain falls from 0.041 on validation to 0.003 on test, while curated data reaches 0.022 but negative 0.019 on test, two non-transfer signatures. Curated external data raises held-out CYP2C9-substrate performance by 0.17 and half-life by 0.08, admitted through a contamination filter rejecting same-source files overlapping 64 to 89 percent of test structures, necessary but not sufficient for transfer. A matched-trial automated machine learning control did not reproduce the agent's code-level model intervention, reaching 0.006 against 0.042, and the pipeline stays competitive with an 84M-parameter pretrained 3D model on the shared training split. The experiments stay within molecular property prediction, but separating discovery from held-out certification is a domain-agnostic lesson for any closed-loop system optimising a proxy for a held-out quantity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。