arXiv:2502.21266cs.DCcs.AI2025-02被引 1

构建联邦云平台,让科研人员轻松用GPU跑机器学习分析

Supporting the development of Machine Learning for fundamental science in a federated Cloud with the AI_INFN platform

  • 用Kubernetes平台统一管理跨地域GPU资源
  • 支持异构分布式环境下的机器学习流程扩展
  • 适合需要共享算力的高能物理等科研团队

机器学习正在改变科学家设计、开发和部署数据密集型软件的方式。然而,ML的采用对计算基础设施提出了新挑战,尤其在硬件加速器的资源配置与调度方面。由意大利国家核物理研究所(INFN)资助的AI_INFN项目旨在通过提供定制化计算资源支持,推动ML技术在INFN研究场景中的应用。该项目基于INFN云平台,采用云原生方案,高效共享硬件加速器,保障研究所各类科研活动的多样性。本文更新了用于支持GPU驱动数据分析工作流开发与扩展的Kubernetes平台的部署进展,该平台可借助interLink提供者实现虚拟kubelet的联邦化部署,支持跨异构分布式资源的灵活调度。

原文摘要 · Abstract (English)

Machine Learning (ML) is driving a revolution in the way scientists design, develop, and deploy data-intensive software. However, the adoption of ML presents new challenges for the computing infrastructure, particularly in terms of provisioning and orchestrating access to hardware accelerators for development, testing, and production. The INFN-funded project AI_INFN ("Artificial Intelligence at INFN") aims at fostering the adoption of ML techniques within INFN use cases by providing support on multiple aspects, including the provision of AI-tailored computing resources. It leverages cloud-native solutions in the context of INFN Cloud, to share hardware accelerators as effectively as possible, ensuring the diversity of the Institute's research activities is not compromised. In this contribution, we provide an update on the commissioning of a Kubernetes platform designed to ease the development of GPU-powered data analysis workflows and their scalability on heterogeneous, distributed computing resources, possibly federated as Virtual Kubelets with the interLink provider.

机器学习联邦云KubernetesGPU加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。