arXiv:2411.00889cs.LGcs.SY2024-11中稿 · the 2024 Workshop …

智能选模型,省电2.5倍还达标

MESS+: Energy-Optimal Inferencing in Language Model Zoos with Service Level Guarantees

  • 按请求实时选最优模型,兼顾能耗与性能
  • 在保证SLA质量前提下,能效提升最高2.5倍
  • 适合关注推理成本与服务质量的系统开发者

开放权重的大语言模型集合让用户可快速将顶尖模型集成到系统中。尽管模型日益丰富,但为特定任务选择最合适的模型仍主要依赖公开排行榜和经验判断,这对推理服务提供商(追求成本效率)和终端用户(追求输出质量)均不理想。在商业场景中,二者常通过服务等级协议(SLA)统一要求。本文提出MESS+,一种面向模型集合的在线随机优化算法,支持按每次推理请求进行能量最优的模型选择。在满足高精度SLA要求的前提下,相比随机选型,能效最高提升2.5倍。

原文摘要 · Abstract (English)

Open-weight large language model (LLM) zoos allow users to quickly integrate state-of-the-art models into systems. Despite increasing availability, selecting the most appropriate model for a given task still largely relies on public benchmark leaderboards and educated guesses. This can be unsatisfactory for both inference service providers and end users, where the providers usually prioritize cost efficiency, while the end users usually prioritize model output quality for their inference requests. In commercial settings, these two priorities are often brought together in Service Level Agreements (SLA). We present MESS+, an online stochastic optimization algorithm for energy-optimal model selection from a model zoo, which works on a per-inference-request basis. For a given SLA that requires high accuracy, we are up to 2.5x more energy efficient with MESS+ than with randomly selecting an LLM from the zoo while maintaining SLA quality constraints.

大模型推理优化能效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。