arXiv:2502.15712cs.NIcs.AI2025-02被引 3

把网卡变成计算单元,降低复杂AI推理的资源开销

GPUs, CPUs, and... NICs: Rethinking the Network's Role in Serving Complex AI Pipelines

  • 将数据处理任务迁移到智能网卡,利用其并行处理能力
  • 实测表明,关键任务可减少70%以上资源占用
  • 适合需要低延迟高吞吐的AI服务部署场景

随着AI应用日益普及,高效管理复杂推理流水线和计算资源的需求愈发突出。当流水线复杂度上升时,分布式部署带来的网络延迟问题日益严重。本文探讨如何将网络从瓶颈转化为助力,通过将资源密集型数据处理任务——这是当前AI流水线复杂性的重要来源——与智能网卡的包处理流水线特性相匹配,实现向SmartNIC的卸载。我们分析了任务卸载面临的挑战与机遇,提出了一项整合网络硬件到AI流水线的研究议程,为系统优化开辟新路径。

原文摘要 · Abstract (English)

The increasing prominence of AI necessitates the deployment of inference platforms for efficient and effective management of AI pipelines and compute resources. As these pipelines grow in complexity, the demand for distributed serving rises and introduces much-dreaded network delays. In this paper, we investigate how the network can instead be a boon to the excessively high resource overheads of AI pipelines. To alleviate these overheads, we discuss how resource-intensive data processing tasks -- a key facet of growing AI pipeline complexity -- are well-matched for the computational characteristics of packet processing pipelines and how they can be offloaded onto SmartNICs. We explore the challenges and opportunities of offloading, and propose a research agenda for integrating network hardware into AI pipelines, unlocking new opportunities for optimization.

AI推理智能网卡资源优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。