arXiv:2601.22001cs.AIcs.AR2026-01

AI代理推理需突破内存瓶颈,靠异构计算实现高效部署

Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference

论文配图:Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference
图 1 · 摘自论文原文
  • 用操作强度和容量足迹分析推理瓶颈,揭示内存墙问题
  • 长上下文缓存使解码高度依赖内存带宽,影响效率
  • 适合关注大模型推理优化与系统级硬件协同设计的读者

AI代理推理正推动以推理为中心的数据中心未来,暴露出超越算力的瓶颈,尤其是内存容量、内存带宽和高速互连。我们提出两个指标——操作强度(OI)和容量足迹(CF),共同解释经典屋顶线模型遗漏的场景,包括内存容量墙。在不同代理工作流(聊天、编程、网页使用、计算机使用)和基础模型选择(GQA/MLA、MoE、量化)下,OI/CF会发生显著变化,长上下文键值缓存使解码阶段高度内存受限。这些观察推动了分离式服务与系统级异构设计:专用预填充和解码加速器、更广泛的扩展网络,以及由光通信实现的解耦计算-内存架构。我们进一步假设,代理-硬件协同设计、单系统内多推理加速器,以及高带宽大容量内存分离,是适应不断演变的OI/CF的基础。这些方向共同为大规模智能体推理的持续效率与能力提供路径。

原文摘要 · Abstract (English)

AI agent inference is driving an inference heavy datacenter future and exposes bottlenecks beyond compute - especially memory capacity, memory bandwidth and high-speed interconnect. We introduce two metrics - Operational Intensity (OI) and Capacity Footprint (CF) - that jointly explain regimes the classic roofline analysis misses, including the memory capacity wall. Across agentic workflows (chat, coding, web use, computer use) and base model choices (GQA/MLA, MoE, quantization), OI/CF can shift dramatically, with long context KV cache making decode highly memory bound. These observations motivate disaggregated serving and system level heterogeneity: specialized prefill and decode accelerators, broader scale up networking, and decoupled compute-memory enabled by optical I/O. We further hypothesize agent-hardware co design, multiple inference accelerators within one system, and high bandwidth, large capacity memory disaggregation as foundations for adaptation to evolving OI/CF. Together, these directions chart a path to sustain efficiency and capability for large scale agentic AI inference.

异构计算推理优化内存瓶颈AI代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。