arXiv:2605.10124cs.NIcs.DC2026-05

GELATO动态调度生成任务,让手机与边缘端协作推理更省电、更快。

GELATO: Generative Entropy- and Lyapunov-based Adaptive Token Offloading for Device-Edge Speculative LLM Inference

论文配图:GELATO: Generative Entropy- and Lyapunov-based Adaptive Token Offloading for Device-Edge Speculative LLM Inference
图 1 · 摘自论文原文
  • 基于熵和李雅普诺夫的自适应调度策略,实时调整每轮生成预算。
  • 在资源受限下提升64.98%吞吐量,能耗降低47.47%,质量无损。
  • 适合移动设备与边缘协同部署的大模型推理场景。

随着移动端大语言模型(LLM)推理的发展,设备-边缘协同推理日益受到关注。其中,推测解码(Speculative Decoding, SD)架构通过轻量级草稿模型快速生成候选词,再由强大目标模型验证,显著提升效率。然而,在资源受限的边缘环境下,如何实现逐令牌资源调度以适配SD仍面临挑战。本文提出一种生成熵与李雅普诺夫结合的自适应令牌卸载框架GELATO,旨在设备-边缘协同的SD系统中,在能量约束下最大化解码吞吐量。具体而言,外层漂移-惩罚环在线决策,建立参考草稿预算,管理长期能效-吞吐权衡;内层熵驱动生成机制实现早期退出,适应每令牌的动态生成不确定性。理论分析给出了GELATO长期吞吐量的严格性能边界。大量实验表明,GELATO实现了全局最优权衡,在资源受限环境下相比现有先进分布式SD架构,吞吐量提升64.98%,能耗降低47.47%,同时保持了大模型解码质量。

原文摘要 · Abstract (English)

The recent growth of on-device Large Language Model (LLM) inference has driven significant interest in device-edge collaborative LLM inference. As a promising architecture, Speculative Decoding (SD) is increasingly adopted where a lightweight draft model rapidly generates candidate tokens to be verified by a powerful target model. However, a fundamental challenge lies in achieving per-token resource scheduling to effectively adapt SD paradigm to resource-constrained edge environment. This paper proposes a Generative Entropy- and Lyapunov-based Adaptive Token Offloading framework, named GELATO, to maximize decoding throughput under energy constraints in a device-edge collaborative SD system. Specifically, an outer drift-plus-penalty loop makes online decisions to establish a reference drafting budget, managing long-term energy-throughput trade-off. Further, a nested entropy-driven generation mechanism executes early exiting to adapt to per-token dynamic generative uncertainty. Theoretical analysis establishes a rigorous performance bound on long-term throughput for GELATO. Extensive evaluations demonstrate that GELATO achieves a globally optimal tradeoff, outperforming state-of-the-art distributed SD architectures by 64.98% in token throughput and reducing energy consumption by 47.47% under resource-constrained environments, while preserving LLM decoding quality.

大模型推理边缘计算动态调度节能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。