arXiv:2503.11246cs.DCcs.AR2025-03被引 2

用四块二手显卡搭建低成本科研集群,助力发展中国家突破算力瓶颈

Cost-effective Deep Learning Infrastructure with NVIDIA GPU

  • 用四块GTX 1650显卡+三台计算节点组建分布式集群
  • 支持大模型训练与科学计算,实测可满足多数科研需求
  • 适合预算有限的高校或研究机构快速部署

深度学习、大数据处理和科研模拟对算力的需求持续增长。尼泊尔等发展中国家常因资源限制难以购置新硬件。本文通过优化现有技术,构建了一个由四块NVIDIA GeForce GTX 1650 GPU组成的计算集群,包含一个主控节点和三个计算节点。主节点部署Anaconda和Slurm实现包管理与任务调度,集成NFS网络存储系统以扩展存储能力。由于集群通过公网域名访问存在安全风险,采用fail2ban防御暴力破解攻击。尽管在设计与实施中面临诸多挑战,本项目证明了利用低成本硬件仍可构建高效算力平台,满足多种高负载科研任务需求。

原文摘要 · Abstract (English)

The growing demand for computational power is driven by advancements in deep learning, the increasing need for big data processing, and the requirements of scientific simulations for academic and research purposes. Developing countries like Nepal often struggle with the resources needed to invest in new and better hardware for these purposes. However, optimizing and building on existing technology can still meet these computing demands effectively. To address these needs, we built a cluster using four NVIDIA GeForce GTX 1650 GPUs. The cluster consists of four nodes: one master node that controls and manages the entire cluster, and three compute nodes dedicated to processing tasks. The master node is equipped with all necessary software for package management, resource scheduling, and deployment, such as Anaconda and Slurm. In addition, a Network File Storage (NFS) system was integrated to provide the additional storage required by the cluster. Given that the cluster is accessible via ssh by a public domain address, which poses significant cybersecurity risks, we implemented fail2ban to mitigate brute force attacks and enhance security. Despite the continuous challenges encountered during the design and implementation process, this project demonstrates how powerful computational clusters can be built to handle resource-intensive tasks in various demanding fields.

算力集群低代码部署科研基建国产替代

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。