arXiv:2510.13724cs.DCcs.AI2025-10被引 7

让科研人员在本地集群用API调用大模型,实现私密高效的推理服务。

FIRST: Federated Inference Resource Scheduling Toolkit for Scientific AI Model Access

  • 通过联邦调度框架,跨多个集群统一管理大模型推理资源。
  • 支持每日生成数十亿文本令牌,低延迟与高吞吐并存。
  • 适合需要私有化部署的科研团队,尤其适用于高性能计算环境。

我们提出联邦推理资源调度工具 FIRST,一个可在分布式高性能计算(HPC)集群间提供推理即服务的框架。FIRST 使研究人员能够通过兼容 OpenAI 的 API,在私有安全环境中访问多种人工智能模型,如大语言模型(LLMs)。依托 Globus Auth 与 Globus Compute,系统支持跨联邦集群分发并行推理任务,可对接多个托管模型。FIRST 支持 vLLM 等多种推理后端,具备自动扩缩容能力,保持“热节点”以实现低延迟执行,并同时提供高吞吐批处理与交互式模式。该框架满足科学工作流中对私密、安全、可扩展推理日益增长的需求,使研究人员可在本地实现每日数亿文本令牌的生成,无需依赖商业云服务。

原文摘要 · Abstract (English)

We present the Federated Inference Resource Scheduling Toolkit (FIRST), a framework enabling Inference-as-a-Service across distributed High-Performance Computing (HPC) clusters. FIRST provides cloud-like access to diverse AI models, like Large Language Models (LLMs), on existing HPC infrastructure. Leveraging Globus Auth and Globus Compute, the system allows researchers to run parallel inference workloads via an OpenAI-compliant API on private, secure environments. This cluster-agnostic API allows requests to be distributed across federated clusters, targeting numerous hosted models. FIRST supports multiple inference backends (e.g., vLLM), auto-scales resources, maintains "hot" nodes for low-latency execution, and offers both high-throughput batch and interactive modes. The framework addresses the growing demand for private, secure, and scalable AI inference in scientific workflows, allowing researchers to generate billions of tokens daily on-premises without relying on commercial cloud infrastructure.

联邦推理大模型科学计算资源调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。