arXiv:2506.05871cs.LGcs.DC2025-06被引 3

BestServe快速估算大模型服务策略的吞吐量,省去繁琐测试。

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures

  • 基于改进的屋顶模型和调度动态模拟推理性能
  • 预测误差小于20%,可在单台普通CPU上分钟级完成优化
  • 适合需要快速部署规划的团队使用

向数百万用户提供大语言模型服务需高效资源分配与并行策略。当前寻找最优策略依赖耗时的试错过程。我们提出BestServe,一个通过估算不同场景下吞吐量来排序服务策略的新框架。该框架支持共置与分离架构,利用基于改进屋顶模型及CPU-GPU调度动态的推理仿真器,在单台标准CPU上仅需几分钟即可确定最优策略,无需昂贵基准测试,预测误差在20%以内。其轻量设计与强可扩展性使其适用于快速部署规划。

原文摘要 · Abstract (English)

Serving large language models (LLMs) to millions of users requires efficient resource allocation and parallelism strategies. It is a labor intensive trial-and-error process to find such a strategy. We present BestServe, a novel framework for ranking serving strategies by estimating goodput under various operating scenarios. Supporting both collocated and disaggregated architectures, BestServe leverages an inference simulator built on an adapted roofline model and CPU-GPU dispatch dynamics. Our framework determines the optimal strategy in minutes on a single standard CPU, eliminating the need for costly benchmarking, while achieving predictions within a $20\%$ error margin. It appeals to be practical for rapid deployment planning because of its lightweight design and strong extensibility.

模型服务吞吐量优化推理框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。