让AI科研可审计:通过隔离测试验证每步决策,提升材料研究可复现性。
Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer

- 以独立轴向搜索+内外层验证,确保每项技术决策可追溯
- 9/10任务外层测试确认最优方案,非绑定决策保留89.3%正确排序
- 适合追求可复现、可验证的自动化材料发现研究者
Auto Research利用语言模型代理在闭环中提出、实施并评估机器学习改进,但通常仅以最终结果评判。终端评分无法揭示哪项技术决策带来收益,也难以区分可复用发现与仅适配开发反馈的调整。本文提出以干预为中心的Auto Research,验证研究决策而非仅最终成果,并使其可靠性可测量。在特征、模型、表示和数据四个维度上分别进行内层五折反馈搜索,各维度最优解冻结后,再通过外层预留矩阵在闭环从未见过的证据上比较所有替代方案。在跨十项Matbench基准的701次代理执行中,外层证据证实了九项任务的选定干预,且保留89.3%的非平局决策顺序。同时,该方法否定了内层反馈支持的聚合表示增益。结果揭示信息依赖层级:仅需组分的任务存在多条优化路径,而结构敏感任务则偏好局部几何特征与互补树集成。后续兼容性测试在不重新搜索或调参的前提下组合已冻结的特征与模型代码,使外层持留平均提升从19.0%增至26.3%。通过验证决策而非仅成果,该设计将自适应搜索转化为可复用证据。
原文摘要 · Abstract (English)
Auto Research uses language-model agents to propose, implement, and evaluate machine-learning changes in a closed loop, but is usually judged by its terminal pipeline. A terminal score cannot reveal which technical decision produced a gain or distinguish a reusable discovery from a change adapted to development feedback. We introduce intervention-centered Auto Research, which validates research decisions rather than only final artifacts and makes their reliability measurable. Feature, Model, Representation, and Data axes are searched independently with inner five-fold feedback. Each axis winner is frozen before an outer-holdout matrix compares all alternatives on evidence the loop never sees. Across 701 agent-executed attempts spanning ten Matbench endpoints, outer evidence confirms the selected intervention on nine of ten endpoints and preserves 89.3\% of non-tied intervention orderings. It also rejects an aggregate Representation gain that inner feedback endorsed. The resulting matrix reveals an information-dependent hierarchy. Composition-only tasks support several routes to improvement, whereas structure-informed tasks favor local geometry features and complementary tree ensembles. A subsequent compatibility test combines already frozen Feature and Model code without further search or tuning and raises mean outer-holdout improvement from 19.0\% to 26.3\%. By validating decisions rather than only artifacts, this design turns adaptive search into reusable evidence wherever agents propose executable alternatives against a fixed evaluator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。