针对混合模型推理,设计芯片岛系统动态调度方案提升效率。
HYDRA: A Heterogeneous Chiplet DSE Framework for Serving Dynamic Hybrid LLM Workloads

- 联合优化芯片布局、带宽分配与运行时调度策略。
- 平均吞吐提升1.55倍,首令牌延迟降低43.7%。
- 适合大规模异构芯片岛部署的LLM服务场景。
混合Transformer-Mamba大语言模型(LLMs)提升了长上下文处理效率,但其异构计算与通信模式给硬件加速带来挑战。基于芯片岛的架构通过集成专用计算与内存单元提供可扩展解法,但静态架构配置与动态运行时策略构成的复杂设计空间难以全面探索。为此,我们提出HYDRA,一个面向异构芯片岛系统上混合LLM服务的完整设计空间探索框架。该框架协同优化芯片岛组成、布局、片间带宽分配、动态批处理与运行时调度。它融合通信感知布局、动态批处理、弹性任务调度,并采用快速马尔可夫性能估计器,以高效准确捕捉多租户运行时动态。在所有工作负载下,相比现有最优基线,平均吞吐提升1.55倍,首令牌延迟降低43.7%,吞吐增益最高达2.3倍。结果表明,架构与运行时策略协同设计对异构芯片岛系统上的大规模LLM服务至关重要。
原文摘要 · Abstract (English)
Hybrid Transformer-Mamba large language models (LLMs) enhance long-context efficiency, but their heterogeneous computation and communication patterns complicate efficient hardware acceleration. Chiplet-based architectures offer a scalable solution by integrating specialized compute and memory units. However, the design space spanning static architectural configurations and dynamic runtime policies is prohibitively large to explore exhaustively. To address this challenge, we present HYDRA, a comprehensive design space exploration framework for hybrid LLM serving on heterogeneous chiplet systems. HYDRA jointly explores chiplet composition, placement, inter-chiplet bandwidth provisioning, dynamic batching, and runtime scheduling. It integrates communication-aware placement, dynamic batching, elastic task scheduling, and a fast Markov-based performance estimator that captures multi-tenant runtime dynamics for efficient and accurate exploration. Across all workloads, HYDRA delivers 1.55x the throughput and 43.7 percent lower time-to-first-token on average, with throughput gains reaching up to 2.3x compared to state-of-the-art baselines. These results highlight that co-designing architecture and runtime policies is critical for efficient large-scale LLM serving on heterogeneous chiplet systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。