arXiv:2601.00823cs.AIcs.IT2026-01

根据推理计算规律智能分配大模型,降低能耗并提升系统稳定性。

Energy-Aware Routing to Large Reasoning Models

  • 基于训练与推理算力规律设计调度策略
  • 找到能耗与波动平衡的最优运行点
  • 适合关注节能部署与系统优化的研究者

大型推理模型(LRMs)的推理能耗因模型类型和推理深度而异。为降低能耗,需合理选择模型并调整运行方式。系统在任务调度至不同LRM时,性能取决于平均能耗供给与随机波动之间的平衡。关键运行点是既不浪费辅助能源也不浪费基础能源的唯一状态:增加基础供给会导致持续过供和能源浪费;减少供给则会依赖辅助能源。在此状态下,性能受限于波动性,因此需通过二阶分析进一步理解——性能由时间、模型及执行选择中波动的吸收能力决定。这一视角强调了考虑方差的路由与调度作为系统设计的核心原则,并为开发节能型模型调度策略提供了理论依据。调度行为在基于训练-计算与推理-计算缩放律的简单策略下被刻画。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) have heterogeneous inference energy costs based on which model is used and how much it reasons. To reduce energy, it is important to choose the right LRM and operate it in the right way. As a result, the performance of systems that dispatch tasks to different individual LRMs depend on the balance between mean energy provisioning and stochastic fluctuations. The critical regime is the unique operating point at which neither auxiliary energy nor baseline energy is systematically wasted. Increasing baseline supply shifts the system toward persistent over-supply and baseline-energy waste, while reducing supply induces persistent reliance on auxiliary energy. Yet in this regime, performance remains volatility-limited and so a second-order characterization provides further insights that we develop. Here, performance is governed by how variability is absorbed across time, models, and execution choices. This perspective highlights variance-aware routing and dispatch as a principled design axis, and provides a theoretical basis for developing energy-aware model routing policies. Routing behavior is characterized when dispatch policies are based on training-compute and inference-compute scaling laws for LRMs.

模型调度能耗优化推理系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。