arXiv:2608.21393cs.AI2026-08

在企业本地部署大模型推理,数据全程不离硬件,实现安全高速的AI响应。

Spyre-Accelerated Retrieval-Augmented Generation on IBM LinuxONE: A Cloud-Native Architecture for Secure, High-Throughput Enterprise AI Inference

  • 全链路部署于IBM LinuxONE,利用Spyre加速生成,数据不出机房。
  • 端到端RAG延迟低于2秒,性能比外平台提升最高20倍。
  • 适合金融、医疗等对数据安全要求严苛的企业用户。

在企业环境中运行大语言模型始终面临现实挑战:数据存于一地,算力在别处,跨系统传输敏感信息带来延迟、安全与合规风险。IBM为LinuxONE及Z系列打造的Spyre加速卡改变了这一局面。本文提出一个六子系统组成的检索增强生成(RAG)架构,完全运行于IBM LinuxONE平台,使用Spyre进行生成式推理,Telum II芯片加速轻量分类任务,通过Red Hat OpenShift实现容器编排。从查询接收、向量检索、提示组装、大模型推理、合规过滤到结果交付,全流程数据均保留在单一LinuxONE系统内,确保敏感数据不离开硬件边界。文中详述各子系统的架构设计,解析Spyre编译与服务栈,说明LinuxONE安全执行技术如何将机密计算保障延伸至AI工作负载,并对比了云GPU与本地部署方案的性能表现。初步分析显示,端到端RAG延迟低于2秒,相较外部平台推理效率提升最高达20倍,同时维持监管行业所需的强加密与可审计性。

原文摘要 · Abstract (English)

Running large language models inside enterprise environments has always bumped up against a practical wall: the data lives in one place, the AI horsepower sits somewhere else, and moving sensitive records between the two creates real headaches around latency, security, and regulatory exposure. IBM's Spyre accelerator PCIe inference card built for LinuxONE and the broader IBM Z family changes that equation. In this paper we lay out a six-subsystem RAG architecture that runs entirely on IBM LinuxONE, using Spyre for generative inference, the Telum II on-chip accelerator for lightweight classification tasks, and Red Hat OpenShift for container orchestration. Every piece of the pipeline from query intake through vector retrieval, prompt assembly, LLM inference, compliance filtering, and response delivery stays within a single LinuxONE system, so sensitive data never has to leave the hardware perimeter. We walk through the design choices behind each subsystem, dig into the Spyre compilation and serving stack, explain how LinuxONE's Secure Execution technology extends confidential-computing guarantees to AI workloads, and benchmark the architecture against cloud-GPU and on-premises alternatives. Early analysis points to end-to-end RAG latencies under two seconds and up to a 20x reduction compared to off-platform inference, all while keeping the strong encryption and auditability posture that regulated industries actually need.

企业AI安全推理RAG硬件加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。