动态调度异构加速器,平衡成本、性能与容错,高效支撑生成式AI推理。
Adaptive Orchestration for Large-Scale Inference on Heterogeneous Accelerator Systems Balancing Cost, Performance, and Resilience
- 基于实时成本与容量信号,自适应分配请求到不同加速器。
- 在稳定扩散模型上实现低延迟、高吞吐,容量不足时自动切换路径。
- 适合需要弹性扩展和成本控制的生成式AI部署场景。
生成式AI工作负载的激增催生了可灵活利用GPU与专用加速器、同时控制运营成本的大规模推理系统需求。本文提出一种硬件无关的控制环,根据实时成本和容量信号,自适应地将请求分配至异构加速器。该方法通过动态切换成本优化与容量优化模式,确保在计算资源波动时仍能保持低延迟与高吞吐,实现昂贵算力的高效利用。在Stable Diffusion模型上的评估表明,该框架持续满足延迟目标,在容量不足时自动重定向流量,并在可能时利用低成本加速器。结果表明,跨软硬件栈的反馈驱动部署策略,有助于组织在加速器资源有限的情况下高效扩展生成式AI工作负载并维持系统韧性。
原文摘要 · Abstract (English)
The surge in generative AI workloads has created a need for scalable inference systems that can flexibly harness both GPUs and specialized accelerators while containing operational costs. This paper proposes a hardware-agnostic control loop that adaptively allocates requests across heterogeneous accelerators based on real-time cost and capacity signals. The approach sustains low latency and high throughput by dynamically shifting between cost-optimized and capacity-optimized modes, ensuring the most efficient use of expensive compute resources under fluctuating availability. Evaluated using the Stable Diffusion model, the framework consistently meets latency targets, automatically redirects traffic during capacity shortfalls, and capitalizes on lower-cost accelerators when possible. These results highlight how a feedback-driven deployment strategy, spanning the entire software and hardware stack, can help organizations efficiently scale generative AI workloads while maintaining resilience in the face of limited accelerator capacity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。