arXiv:2410.14257cs.LGcs.AI2024-10被引 10

提出新指标框架,更好衡量大模型服务的用户体验与系统性能平衡。

Revisiting Service Level Objectives and System Level Metrics in Large Language Model Serving

  • 设计符合用户实际体验的新SLO,避免人为延迟提升指标的误导。
  • 提出平滑吞吐量(smooth goodput)统一评估请求响应与系统效率。
  • 适合关注大模型服务优化、用户体验与系统调度的研究者参考。

用户体验是大语言模型(LLM)服务系统的关键考量因素,其中面向单个请求的服务水平目标(SLO)和面向整体系统性能的系统级指标(SLMs)是两大核心度量标准。然而,现有指标存在两个显著问题:1)人为延迟部分令牌输出可改善SLO;2)主动放弃不符合SLO的请求可提升SLMs,二者均违背直觉。本文重新审视LLM服务中的SLO与SLMs,提出一个与用户真实体验对齐的新SLO。基于此,提出名为平滑吞吐量(smooth goodput)的综合指标框架,整合SLO与SLMs,反映LLM服务中用户体验的本质。通过该统一框架,我们在多种工作负载下重新评估了不同LLM服务系统的性能。实验结果表明,该框架能更全面地刻画令牌交付与请求处理情况,有效识别不同服务策略下用户体验与系统性能的最佳平衡点。

原文摘要 · Abstract (English)

User experience is a critical factor Large Language Model (LLM) serving systems must consider, where service level objectives (SLOs) considering the experience of individual requests and system level metrics (SLMs) considering the overall system performance are two key performance measures. However, we observe two notable issues in existing metrics: 1) manually delaying the delivery of some tokens can improve SLOs, and 2) actively abandoning requests that do not meet SLOs can improve SLMs, both of which are counterintuitive. In this paper, we revisit SLOs and SLMs in LLM serving, and propose a new SLO that aligns with user experience. Based on the SLO, we propose a comprehensive metric framework called smooth goodput, which integrates SLOs and SLMs to reflect the nature of user experience in LLM serving. Through this unified framework, we reassess the performance of different LLM serving systems under multiple workloads. Evaluation results show that our metric framework provides a more comprehensive view of token delivery and request processing, and effectively captures the optimal point of user experience and system performance with different serving strategies.

大模型服务性能评估用户体验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。