arXiv:2608.25543eess.SYcs.AI2026-08被引 1

提升边缘端大模型推理吞吐量,确保服务达标

Goodput Maximization for Large Language Model Edge Inference: A Two-Phase Maskable PPO Approach

论文配图:Goodput Maximization for Large Language Model Edge Inference: A Two-Phase Maskable PPO Approach
图 1 · 摘自论文原文
  • 分两阶段优化任务卸载与带宽分配,用掩码减少无效动作
  • 相比基准方法,系统奖励提升33.3%至87.5%,好吞吐最高
  • 适合无线边缘网络中大模型推理的高并发场景

本文提出一种新型两阶段可掩码近端策略优化(TP-MPPO)算法,旨在无线边缘网络中大语言模型(LLM)推理服务下,最大化系统好吞吐量(goodput),同时严格满足服务等级目标(SLO)。第一阶段通过带动作掩码的近端策略优化(MPPO)优化任务卸载决策,有效避免无效动作探索,缩小动作空间;第二阶段推导上行带宽分配的闭式解,设计贪心算法完成下行带宽分配,为下一回合的MPPO提供即时奖励。两阶段交替进行直至收敛。仿真结果表明,与基准方法相比,TP-MPPO可使系统奖励提升33.3%–87.5%,并实现最高好吞吐量。

原文摘要 · Abstract (English)

This paper presents a novel two-phase maskable proximal policy optimization (TP-MPPO) algorithm, which maximizes the system goodput counting request throughput with strict service level objective (SLO) compliance for large language model (LLM) inference services in wireless edge networks. In the first phase of TP-MPPO, we optimize the task offloading decisions by MPPO with action masking mechanism, effectively avoiding exploring invalid actions and reducing the action space. In the second phase, closed-form solutions are derived for uplink bandwidth allocation; a greedy algorithm is designed for downlink bandwidth allocation to provide immediate rewards for the MPPO in the next round. The two stages alternate till convergence. Simulation results demonstrate that TP-MPPO can improve the system reward by 33.3%--87.5% compared to its benchmarks and achieve the highest goodput.

边缘计算大模型推理强化学习资源调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。