arXiv:2503.13343cs.DCcs.AI2025-03被引 1

构建可扩展运行时架构,实现HPC与ML协同调度。

Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications

  • 基于服务化执行扩展RADICAL-Pilot,支持跨平台混合工作流。
  • 在本地与远程资源上并行运行多个ML模型,开销极低。
  • 适合需要大规模科学计算与机器学习融合的研究者使用。

结合传统高性能计算(HPC)与新型机器学习(ML)方法的混合工作流正在重塑科学计算。本文提出并实现了可扩展的运行时系统,通过服务化执行扩展RADICAL-Pilot,支持AI融入HPC的工作流。该系统实现了分布式机器学习能力、高效的资源管理,并在本地与远程平台间实现无缝的HPC/ML耦合。初步实验表明,该方法可在本地与远程HPC/云资源上高效并行执行多个机器学习模型,且架构开销极小。这为原型化三个典型数据驱动工作流应用并将其在领导级超算平台上规模化执行奠定了基础。

原文摘要 · Abstract (English)

Hybrid workflows combining traditional HPC and novel ML methodologies are transforming scientific computing. This paper presents the architecture and implementation of a scalable runtime system that extends RADICAL-Pilot with service-based execution to support AI-out-HPC workflows. Our runtime system enables distributed ML capabilities, efficient resource management, and seamless HPC/ML coupling across local and remote platforms. Preliminary experimental results show that our approach manages concurrent execution of ML models across local and remote HPC/cloud resources with minimal architectural overheads. This lays the foundation for prototyping three representative data-driven workflow applications and executing them at scale on leadership-class HPC platforms.

HPC机器学习工作流可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。