通过联合优化提示压缩与无线功率,提升大模型服务效率。
JPPO: Joint Power and Prompt Optimization for Accelerated Large Language Model Services
- 用小模型在用户端压缩提示,减少传输数据量
- 深度强化学习协调压缩率与发射功率,降低17%响应时间
- 适合资源受限的移动大模型应用
大语言模型(LLMs)在各类任务中表现出色,正被广泛部署于无线网络以支持多样化的用户服务。然而,日益增长的长提示设置带来了巨大的计算资源需求和通信负载。为此,我们提出联合功率与提示优化(JPPO)框架,结合小语言模型(SLM)的提示压缩与无线功率分配优化。在用户设备部署SLM进行提示压缩,并采用深度强化学习联合优化压缩比与传输功率,有效平衡服务质量和资源效率。实验结果表明,该框架在保持高服务保真度和低误码率的同时,优化了无线大模型服务的功耗。系统响应时间平均降低约17%,优化效果随原始提示长度变化而异。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities in various tasks, leading to their increasing deployment in wireless networks for a wide variety of user services. However, the growing longer prompt setting highlights the crucial issue of computational resource demands and huge communication load. To address this challenge, we propose Joint Power and Prompt Optimization (JPPO), a framework that combines Small Language Model (SLM)-based prompt compression with wireless power allocation optimization. By deploying SLM at user devices for prompt compression and employing Deep Reinforcement Learning for joint optimization of compression ratio and transmission power, JPPO effectively balances service quality with resource efficiency. Experimental results demonstrate that our framework achieves high service fidelity and low bit error rates while optimizing power usage in wireless LLM services. The system reduces response time by about 17%, with the improvement varying based on the length of the original prompt.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。