arXiv:2606.25879cs.DCcs.AI2026-06

用AI助手在FABRIC测试平台复现多个科研实验,效率提升4到6倍。

AI-Assisted Computational Reproducibility on the FABRIC Testbed

  • 用LLM辅助配置环境、改代码、调试,降低复现门槛。
  • 复现实验支持原研究结论,但分析阶段需人工介入流程设计。
  • 适合想高效复现实验的科研人员和测试平台开发者。

计算可复现性虽是科学研究的核心,却仍面临挑战。本文展示如何利用国际FABRIC测试平台,结合大语言模型编码助手LoomAI,简化跨领域的实验复现。我们在FABRIC上复现了三个案例:BBR族拥塞控制评估、仅使用CPU的MPI集群上LAMMPS分子动力学扩展性基准测试,以及应激蛋白稳态基因组分析流程。复现不仅关注数值结果一致,更验证是否支持原研究的科学结论。AI助手在环境搭建、代码适配和调试中表现良好,但在缺乏明确工作流的分析阶段表现不足,需人工指导执行顺序与数据依赖关系。整体上,AI辅助流程使复现工作量减少约4至6倍。最后提出改进研究测试平台中AI辅助可复现性的实践建议。

原文摘要 · Abstract (English)

Computational reproducibility remains difficult despite being central to scientific research. In this paper, we show how the international FABRIC testbed, combined with large language model (LLM) coding assistants through LoomAI, can simplify reproducing published experiments across multiple domains. We reproduced three case studies on FABRIC, covering BBR-family congestion-control evaluations, LAMMPS molecular dynamics scaling benchmarks on a CPU-only MPI cluster, and stress protein homeostasis genomics pipelines. Rather than focusing only on matching numerical outputs, we evaluate whether the reproduced experiments support the same scientific conclusions as the original studies. The AI assistant was effective in setting up the environment, adapting code, and debugging, but struggled with the analysis stages that lacked clearly defined workflows, which required human guidance to establish execution order and data dependencies. Across the case studies, the AI-assisted workflow reduced reproduction effort by roughly 4--6 times. We conclude with practical recommendations for improving AI-assisted reproducibility on research testbeds.

可复现性AI辅助测试平台实验复现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。