arXiv:2603.27624cs.ARcs.AI2026-03

通过芯片组动态调度专家流,加速边缘设备低批量MoE推理。

Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling

  • 设计多芯片组架构下的细粒度专家流调度策略。
  • 相比现有方法提速1.22至2.00倍,节省78.8%片上内存。
  • 适合资源受限的边缘AI部署,尤其适用于低批量MoE场景。

混合专家(MoE)在边缘AI低批量推理中具有潜力,但设备端部署常受限于片上内存和严重的负载不均衡;普遍采用的卸载机制进一步引发片外内存访问瓶颈。此外,MoE稀疏性和动态门控使分布式策略趋向更细粒度,并引入运行时调度需求。近期高带宽的芯粒间互连为多芯粒系统解决负载不均衡与卸载瓶颈提供了新可能。本文提出全分片专家数据并行(FSE-DP),专为多芯粒加速器上的低批量MoE推理设计的并行范式。FSE-DP通过跨高带宽芯粒间链路协调细粒度、互补的专家流沿动态轨迹运行,实现自适应计算-通信重叠与负载均衡。数据流复杂性通过一组极简、硬件友好的虚拟化规则和轻量级调度算法得以控制。该方法相较最先进基线提升1.22至2.00倍速度,节省高达78.8%的片上内存。

原文摘要 · Abstract (English)

Mixture-of-Experts is a promising approach for edge AI with low-batch inference. Yet, on-device deployments often face limited on-chip memory and severe workload imbalance; the prevalent use of offloading further incurs off-chip memory access bottlenecks. Moreover, MoE sparsity and dynamic gating shift distributed strategies toward much finer granularity and introduce runtime scheduling considerations. Recently, high die-to-die bandwidth chiplet interconnects have created new opportunities for multi-chiplet systems to address workload imbalance and offloading bottlenecks with fine-grained scheduling. In this paper, we propose Fully Sharded Expert Data Parallelism, a parallelization paradigm specifically architected for low-batch MoE inference on multi-chiplet accelerators. FSE-DP attains adaptive computation-communication overlap and balanced load by orchestrating fine-grained, complementary expert streams along dynamic trajectories across high-bandwidth D2D links. The attendant dataflow complexity is tamed by a minimal, hardware-amenable set of virtualization rules and a lightweight scheduling algorithm. Our approach achieves 1.22 to 2.00 times speedup over state-of-the-art baselines and saves up to 78.8 percent on-chip memory.

MoE边缘计算芯片组推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。