arXiv:2511.01695cs.LGeess.SP2025-11被引 2

通过智能调度提升移动端大模型推理速度

Collaborative Large Language Model Inference via Resource-Aware Parallel Speculative Decoding

  • 用多智能体强化学习统一优化用户与资源分配
  • 端到端延迟平均降低23.7%,最高达28.0%
  • 适合边缘计算中资源受限的实时大模型应用

移动端大模型推理需求增长,亟需高效移动边缘计算(MEC)方案,尤其在资源受限场景。推测解码通过将生成任务分摊至移动设备上的轻量级草稿模型与边缘服务器的强大目标模型,提供可行路径,但存在通信开销和异步延迟问题。本文首次提出统一框架,联合优化用户关联与资源分配(UARA),以支持高效的并行推测解码。采用多智能体深度强化学习求解UARA问题。基于Sionna模拟器开展真实场景实验,结果表明:本方法在不牺牲推理准确率的前提下,端到端延迟平均降低23.7%,最高达28.0%,可实现可扩展、低延迟的大模型服务。

原文摘要 · Abstract (English)

The growing demand for on-device large language model (LLM) inference highlights the need for efficient mobile edge computing (MEC) solutions, especially in resource-constrained settings. Speculative decoding offers a promising solution by partitioning token generation between a lightweight draft model on mobile devices and a powerful target model on edge servers, but suffers from communication overhead and asynchronous delays. This paper is the first to propose a unified framework that jointly optimizes user association and resource allocation (UARA) to support efficient parallel speculative decoding. We solve the UARA problem using a multi-agent deep reinforcement learning algorithm. To evaluate our approach under realistic conditions, we conduct experiments using the Sionna simulator. Results show that our method achieves up to 28.0% and an average of 23.7% reduction in end-to-end latency without compromising inference accuracy, enabling scalable and low-latency LLM services in MEC systems.

大模型推理边缘计算强化学习推测解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。