自动提炼开源代码库,提升科学编程代理的实验生成能力。
CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents
- 从250个材料科学仓库中自动提取可运行的领域代码片段。
- 生成的代码库使科学代理实验准确率和完整性提升74%。
- 适合开发科学计算自动化系统的研究者使用。
自动化科学发现(ASD)系统可自动生成并运行基于代码的实验,但其能力受限于仅凭参数知识生成代码的能力。当前系统要么仅修改少量手工编写的实验示例,要么完全依赖参数知识,导致质量和覆盖范围受限。我们提出CodeDistiller,一个将大规模科学开源仓库自动提炼为经过验证的领域专用代码库的系统,使ASD代理无需人工干预即可扩展能力。在250个材料科学仓库上结合自动与领域专家评估,最佳模型能在74%的仓库中生成有效代码;下游评估显示,使用CodeDistiller生成的代码库增强后的ASD代理,比仅使用通用材料科学代码的代理产生更准确、完整且科学合理的实验。此外,我们在A/B测试中比较了大模型作为裁判的评分与领域专家评分,发现中等一致性,表明低成本代理指标可能适用于大规模评估科学发现系统。
原文摘要 · Abstract (English)
Automated Scientific Discovery (ASD) systems can help automatically generate and run code-based experiments, but their capabilities are limited by the code they can reliably generate from parametric knowledge alone. As a result, current systems either mutate a small number of manually-crafted experiment examples, or operate solely from parametric knowledge, limiting quality and reach. We introduce CodeDistiller, a system that automatically distills large collections of scientific Github repositories into a vetted library of working domain-specific code examples, allowing ASD agents to expand their capabilities without manual effort. Using a combination of automatic and domain-expert evaluation on 250 materials science repositories, we find the best model is capable of producing functional examples for 74% of repositories, while our downstream evaluation shows an ASD agent augmented with a CodeDistiller generated library produces more accurate, complete, and scientifically sound experiments than an agent with only general materials-science code examples. We also evaluate LLM-as-a-judge ratings against domain-expert ratings in an A/B testing paradigm, finding moderate agreement and suggesting that inexpensive proxy metrics may be feasible for evaluating scientific discovery systems at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。