arXiv:2606.07491cs.DCcs.AI2026-06

教科研人员设计高效可复现的智能超算工作流。

Twelve quick tips for designing AI-driven HPC workflows

论文配图:Twelve quick tips for designing AI-driven HPC workflows
图 1 · 摘自论文原文
  • 用容器化、任务数组等技术应对AI工作流的动态性挑战。
  • 解决小文件频繁I/O和资源异构问题,提升计算效率。
  • 适合从事计算生物学等高负载科研的团队参考。

高性能计算(HPC)集群仍是大规模科学计算的核心,传统上执行确定性、线性优化的流水线。然而,人工智能(AI)与基础模型在科研中的广泛集成,引入了迭代、数据驱动且概率性的新计算范式,带来数据引力、异构资源管理及复杂工作流编排等新挑战。本文提出十二项实用建议,帮助研究人员设计高效、可扩展且可复现的AI驱动型HPC工作流。通过解决容器化环境移植、作业数组策略部署、显式反馈机制构建、小文件I/O优化等关键系统瓶颈,本文提供从刚性流水线向自适应智能计算环境过渡的框架。这些架构原则广泛适用于分布式环境,尤其契合现代计算生物学对资源密集型吞吐量的需求。

原文摘要 · Abstract (English)

High-performance computing (HPC) clusters remain the backbone of large-scale scientific computation, traditionally executing deterministic, linear pipelines optimised for predictable performance. However, the pervasive integration of artificial intelligence (AI) and foundation models into scientific research has introduced a fundamentally new computational paradigm. AI-driven workflows are characteristically iterative, data-driven, and probabilistic, introducing unique challenges regarding data gravity, heterogeneous resource management, and complex workflow orchestration. This guide provides twelve practical tips designed to help researchers design efficient, scalable, and reproducible AI-driven HPC workflows. By addressing critical system-level bottlenecks - such as containerisation for environment portability, strategic deployment of job arrays, explicit feedback loop mechanics, and I/O optimisation for small files - this article offers a framework for transitioning from rigid execution pipelines to adaptive, intelligent computational environments. While these architectural principles are broadly applicable across distributed environments, they are particularly tailored to the resource-intensive throughput demands of modern computational biology.

AI工作流超算计算生物学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。