让人类在超算中异步参与AI训练,不阻塞计算任务
A Workflow-Oriented Framework for Asynchronous Human-AI Collaboration in Hybrid and Compute-Intensive HPC Environments

- 在关键节点暂停流程等待人工输入,计算仍后台运行
- 支持混合环境(超算/本地/云)和SLURM调度系统
- 适合需要人工判断的高安全场景,提升可监督性
在国防与安全等高风险领域,人类在训练和部署AI系统中至关重要。然而,由于高性能计算(HPC)环境的计算强度和资源限制,实时交互难以实现。本文提出一种面向工作流的异步人机协作框架,适用于混合基础设施,包括HPC集群、本地设备和云平台。该框架可在预设检查点暂停工作流以获取人工输入,同时保持底层计算任务持续运行,避免资源闲置并实现非阻塞式监管。其支持基于SLURM的调度、容器化与原生任务,并针对需人工判断与灵活性的场景进行了定制。我们在MareNostrum 5系统上展示了模型训练中的应用,验证了该框架在可移植性、效率和操作监督方面的优势。
原文摘要 · Abstract (English)
Human involvement is critical in training and deploying AI systems in high-stakes defence and security contexts. However, real-time interaction is impractical in HPC environments due to compute intensity and resource constraints. We present a workflow framework that enables asynchronous human-AI collaboration across hybrid infrastructures, including HPC clusters, local machines, and cloud platforms. Workflows can pause at defined checkpoints for human input without halting underlying compute jobs, preventing idle resources and enabling non-blocking supervision. The framework supports interaction with SLURM-based scheduling, containerized and native tasks, and is customized for scenarios requiring human judgment and adaptability. We demonstrate its application in model training on systems like MareNostrum 5, highlighting benefits in portability, efficiency, and oversight in operational AI workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。