arXiv:2608.13573cs.AI2026-08

基于一年真实服务数据,揭示大模型推理负载的演化规律。

A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing

论文配图:A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
图 1 · 摘自论文原文
  • 分析一年生产环境中的多模型、多用户真实请求轨迹。
  • 发现负载随时间变化及长尾模型占比显著,与聚合统计结果不同。
  • 开源完整数据集,支持未来真实场景下的服务系统研究。

大型语言模型(LLM)服务已成为关键云工作负载,真实负载轨迹对系统设计和基准测试至关重要。然而,现有研究在规模和范围上仍受限,常仅观察短期数据,难以全面反映用户在实际生产环境中与模型的交互行为。本文通过对Chutes平台一整年生产级负载轨迹进行全局表征与纵向分析,首次揭示了多模型、多用户场景下负载的演化过程与用户-模型交互结构。相比以往研究,本工作涵盖主流与长尾模型,从聚合、时间、模型和用户四个层面深入分析,揭示了隐藏在平均值背后的动态特征。为促进后续研究,论文将公开完整的年度数据集,使研究者可基于真实流量开展实验,避免依赖抽样或合成数据。

原文摘要 · Abstract (English)

Large Language Model (LLM) serving has become a critical cloud workload, and realistic traces are essential for motivating and benchmarking serving systems. However, existing LLM serving workload studies remain limited in scale and scope. They often observe short time periods and provide limited visibility into how users interact with models in production. As a result, they do not fully capture how LLM serving workloads evolve over time or how user-model interactions shape production traffic. In this work, we further the understanding of real-world LLM serving workloads through both a global characterization and a longitudinal study of a one-year production trace from Chutes. Unlike prior studies, our trace captures full production behavior across many models and users, including both popular and long-tail models. We analyze the workload from aggregate, temporal, model-level, and user-level perspectives, revealing workload evolution and user-model structure that are typically hidden behind aggregate views. To support future research, we will release the full one-year trace with the paper, enabling downstream studies of production behavior without relying on sampled or synthetically generated workloads.

LLM服务负载分析生产数据长期追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。