arXiv:2601.03513cs.SEcs.AI2026-01被引 2

一天内自动化部署五万多个科学工具,解决开源软件难用难题

Deploy-Master: Automating the Deployment of 50,000+ Agent-Ready Scientific Tools in One Day

  • 构建智能工作流,自动发现并验证超50万仓库中的可运行工具
  • 成功构建50,112个可复现的容器化运行环境,实现一键调用
  • 揭示大规模部署中的失败模式与成本特征,推动科研自动化

开源科学软件虽多,但多数难以编译、配置和复用,长期维持小作坊式科研计算模式。这一部署瓶颈限制了可复现性、大规模评估以及科学工具在AI for Science(AI4S)和智能体工作流中的实际集成。我们提出Deploy-Master,一个一站式智能工作流,实现大规模工具发现、构建规格推断、基于执行的验证与发布。基于涵盖90多个科学工程领域的分类体系,发现阶段从超过50万公开仓库出发,经许可与质量筛选,最终获得52,550个可执行工具候选。Deploy-Master将异构开源仓库转化为基于实际执行而非文档声明的可运行容器化能力。单日内完成52,550次构建尝试,成功创建50,112个科学工具的可复现运行环境。每个成功工具通过最小可执行命令验证,并注册至SciencePedia供搜索与复用,支持人工直接使用或代理调用。除交付可用工具外,我们还报告了覆盖5万工具级别的部署轨迹,揭示吞吐量、成本分布、失败面及规格不确定性等规模效应下的新现象,解释科学软件难于落地的根本原因,并呼吁建立共享、可观测的执行底座,作为可扩展AI4S与智能体科研的基础。

原文摘要 · Abstract (English)

Open-source scientific software is abundant, yet most tools remain difficult to compile, configure, and reuse, sustaining a small-workshop mode of scientific computing. This deployment bottleneck limits reproducibility, large-scale evaluation, and the practical integration of scientific tools into modern AI-for-Science (AI4S) and agentic workflows. We present Deploy-Master, a one-stop agentic workflow for large-scale tool discovery, build specification inference, execution-based validation, and publication. Guided by a taxonomy spanning 90+ scientific and engineering domains, our discovery stage starts from a recall-oriented pool of over 500,000 public repositories and progressively filters it to 52,550 executable tool candidates under license- and quality-aware criteria. Deploy-Master transforms heterogeneous open-source repositories into runnable, containerized capabilities grounded in execution rather than documentation claims. In a single day, we performed 52,550 build attempts and constructed reproducible runtime environments for 50,112 scientific tools. Each successful tool is validated by a minimal executable command and registered in SciencePedia for search and reuse, enabling direct human use and optional agent-based invocation. Beyond delivering runnable tools, we report a deployment trace at the scale of 50,000 tools, characterizing throughput, cost profiles, failure surfaces, and specification uncertainty that become visible only at scale. These results explain why scientific software remains difficult to operationalize and motivate shared, observable execution substrates as a foundation for scalable AI4S and agentic science.

科学计算自动化部署AI for Science工具链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。