构建全球多模型大模型服务细粒度数据集,揭示真实负载动态差异。
FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads

- 从全球商用平台采集真实服务负载,实现跨模型、跨任务的细粒度建模
- 发现不同模型架构与任务类型存在根本不同的请求波动模式
- 生成可配置的混合负载,用于评估多模型部署中的调度与容量规划
大型语言模型(LLMs)正日益作为持续在线的服务部署,高效的大模型服务成为关键系统挑战。在需求波动下实现低延迟与高吞吐,需深入理解真实服务负载,但现有研究多依赖代理追踪或粗粒度表征,难以捕捉现代多模型平台的异构性。本文提出FineServe,一个来自全球商业市场的在野多模型大模型服务工作负载数据集,支持对异构模型与任务间真实服务动态的细粒度刻画。基于FineServe,我们全面分析了请求到达特性与令牌行为,揭示了不同模型架构、规模及任务意图下的根本性波动模式差异。在此基础上,我们开发了FineServe工作负载生成器,可将细粒度的模型感知负载组合成可配置混合负载,适用于多模型服务系统的基准测试。通过暴露这些细粒度负载动态,FineServe为路由、调度与容量规划策略的评估提供了真实基础。数据集已开源:https://github.com/hihiztc1/FineServe。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed as always-on online services, making efficient LLM serving a critical systems challenge. Achieving low latency and high throughput under volatile demand requires deep understanding of real-world serving workloads, yet existing studies often rely on proxy traces or coarse-grained characterizations that fail to capture the heterogeneity of modern multi-model LLM platforms. We present FineServe, an in-the-wild, multi-model LLM serving workload dataset collected from a global commercial marketplace, enabling fine-grained characterization of real-world serving dynamics across heterogeneous models and tasks. Leveraging FineServe, we conduct a comprehensive analysis of arrival dynamics and token behavior, revealing fundamentally different fluctuation regimes across model architectures, scales and task intents. Building on these insights, we develop the FineServe workload generator, which composes fine-grained model-aware workloads into configurable mixtures tailored for benchmarking multi-model serving platforms. By exposing these fine-grained workload dynamics, FineServe provides a realistic foundation for evaluating routing, scheduling, and capacity-planning strategies in LLM serving systems. FineServe is available at https://github.com/hihiztc1/FineServe.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。