让手机边缘设备高效运行大模型推理,兼顾速度与资源消耗。
Resource-Aware LLM Reasoning for Mobile Edge General Intelligence
- 通过动态调整推理深度和专家网络,智能分配计算资源。
- 推理延迟增加不足1秒,准确率与响应率均达90%以上。
- 适合在算力有限的移动边缘设备上部署复杂大模型任务。
大语言模型(LLM)的发展推动了具备强大推理与自主决策能力的智能体人工智能兴起。结合边缘计算,催生了移动边缘通用智能(MEGI),可在网络边缘实现实时、隐私保护的推理。然而,将基于LLM的智能体推理部署于MEGI环境面临显著挑战:推理计算开销大,而边缘设备资源有限。为此,我们提出一种联合优化框架,实现MEGI中高效LLM推理部署。首先系统梳理增强方法,筛选适用于边缘适配的机制;随后设计分布式架构,融合自适应思维链提示与可扩展的分布式MoE结构。关键创新在于将推理深度建模为动态网络资源变量,与专家激活和传输功率联合优化。该机制可根据任务需求和设备能力动态调节专家网络与推理复杂度。在移动边缘环境中的实验表明,该框架有效平衡推理质量与资源效率:额外推理时间低于1秒,准确率与延迟满足率均达90%,验证了在资源受限的MEGI系统中部署复杂LLM推理的可行性。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) has enabled an emergence of agentic artificial intelligence (AI) with powerful reasoning and autonomous decision-making capabilities. This integration with edge computing has led to the development of Mobile Edge General Intelligence (MEGI), which brings real-time, privacy-preserving reasoning to the network edge. However, deploying LLM-based agentic AI reasoning in MEGI environments poses significant challenges due to the high computational demands of reasoning and the limited resources of edge devices. To address these challenges, we propose a joint optimization framework for efficient LLM reasoning deployment in MEGI. First, we systematically review enhancement methods to identify mechanisms suitable for edge adaptation. Subsequently, we present a distributed framework that synergizes reasoning enhancement via adaptive CoT prompting with scalable deployment through a distributed MoE architecture. An important innovation of this approach involves modeling reasoning depth as a dynamic network resource variable, which is optimized jointly with expert activation and transmission power. This mechanism allows the system to dynamically regulate expert networks and reasoning complexity according to task requirements and device capabilities. Experimental evaluations in mobile edge environments demonstrate that the proposed framework effectively balances reasoning quality and resource efficiency. The results show that with less than one second of additional inference time, both accuracy and latency satisfaction rate can reach 90\%, validating the practical viability of deploying sophisticated LLM reasoning in resource-constrained MEGI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。