arXiv:2507.01880cs.DCcs.LG2025-07被引 4

为提升超算系统对机器学习的支持,提出一套软硬件协同增强方案。

Evolving HPC services to enable ML workloads on HPE Cray EX

  • 构建适配机器学习的用户环境与服务架构,兼顾性能与灵活性
  • 部署性能监测工具与可观测性系统,支持大规模训练过程监控
  • 专为机器学习设计存储与节点验证机制,适合科研团队快速上手

Alps研究基础设施基于GH200技术规模部署,拥有10,752块GPU,为人工智能与机器学习研究者提供强大算力支持。尽管Alps服务于广泛的科学领域,传统HPC服务难以满足机器学习社区的动态需求。本文针对瑞士AI社区自2023年早期访问阶段以来发现的关键挑战与缺口,提出多项技术改进:包括面向机器学习的用户环境,平衡性能与灵活性;用于开发阶段快速评估模型性能的工具;支持大规模训练任务可观测性的数据产品与监控能力;简化节点资源可用性验证的工具;支持训练、推理等多类型工作负载的服务平面架构;以及针对机器学习特点定制的存储基础设施。这些改进旨在提升机器学习任务在超算系统上的执行效率、系统易用性与鲁棒性,并更好地契合机器学习社区需求。同时讨论了当前安全策略。最后,将这些举措置于超算服务对象变化的大背景中进行审视。

原文摘要 · Abstract (English)

The Alps Research Infrastructure leverages GH200 technology at scale, featuring 10,752 GPUs. Accessing Alps provides a significant computational advantage for researchers in Artificial Intelligence (AI) and Machine Learning (ML). While Alps serves a broad range of scientific communities, traditional HPC services alone are not sufficient to meet the dynamic needs of the ML community. This paper presents an initial investigation into extending HPC service capabilities to better support ML workloads. We identify key challenges and gaps we have observed since the early-access phase (2023) of Alps by the Swiss AI community and propose several technological enhancements. These include a user environment designed to facilitate the adoption of HPC for ML workloads, balancing performance with flexibility; a utility for rapid performance screening of ML applications during development; observability capabilities and data products for inspecting ongoing large-scale ML workloads; a utility to simplify the vetting of allocated nodes for compute readiness; a service plane infrastructure to deploy various types of workloads, including support and inference services; and a storage infrastructure tailored to the specific needs of ML workloads. These enhancements aim to facilitate the execution of ML workloads on HPC systems, increase system usability and resilience, and better align with the needs of the ML community. We also discuss our current approach to security aspects. This paper concludes by placing these proposals in the broader context of changes in the communities served by HPC infrastructure like ours.

机器学习超算系统优化基础设施

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。