arXiv:2409.16537cs.LG2024-09

提出ERA算法,平衡边缘智能中的延迟、体验和资源消耗。

A QoE-Aware Split Inference Accelerating Algorithm for NOMA-based Edge Intelligence

  • 综合考虑延迟、用户体验和资源消耗,优化模型拆分与资源分配。
  • 通过梯度下降实现三者最优权衡,实验显示性能显著优于旧方法。
  • 专为边缘智能设计,适合关注用户体验的部署场景。

尽管人工智能已广泛使用并深刻改变生活,但在资源受限的边缘设备上直接部署大型AI模型并不合适。因此,模型拆分推理被提出以提升边缘智能性能:将AI模型分割为多个子模型,将计算密集型部分无线卸载至边缘服务器,以降低资源需求和推理延迟。然而,以往研究主要聚焦于系统QoS优化,忽略了用户体验(QoE)这一关键因素。尽管QoE在边缘计算(EC)中已有研究,但任务卸载与边缘智能(EI)中的拆分推理存在差异,且现有算法未解决EI中的特定QoE问题。为此,本文提出一种有效资源分配算法——ERA,用于加速边缘智能中的拆分推理,并在推理延迟、用户感知质量(QoE)与资源消耗之间实现权衡。具体而言,ERA同时考虑资源消耗、QoE与推理延迟,寻找最优模型拆分与资源分配策略。由于最小延迟、最低资源消耗与最高QoE无法同时满足,采用基于梯度下降的算法寻找最优折衷解。此外,提出循环迭代梯度下降方法,降低因参数离散化带来的算法复杂度。还分析了算法的收敛性、复杂度与近似误差。实验结果表明,ERA性能显著优于先前方法。

原文摘要 · Abstract (English)

Even the AI has been widely used and significantly changed our life, deploying the large AI models on resource limited edge devices directly is not appropriate. Thus, the model split inference is proposed to improve the performance of edge intelligence, in which the AI model is divided into different sub models and the resource-intensive sub model is offloaded to edge server wirelessly for reducing resource requirements and inference latency. However, the previous works mainly concentrate on improving and optimizing the system QoS, ignore the effect of QoE which is another critical item for the users except for QoS. Even the QoE has been widely learned in EC, considering the differences between task offloading in EC and split inference in EI, and the specific issues in QoE which are still not addressed in EC and EI, these algorithms cannot work effectively in edge split inference scenarios. Thus, an effective resource allocation algorithm is proposed in this paper, for accelerating split inference in EI and achieving the tradeoff between inference delay, QoE, and resource consumption, abbreviated as ERA. Specifically, the ERA takes the resource consumption, QoE, and inference latency into account to find the optimal model split strategy and resource allocation strategy. Since the minimum inference delay and resource consumption, and maximum QoE cannot be satisfied simultaneously, the gradient descent based algorithm is adopted to find the optimal tradeoff between them. Moreover, the loop iteration GD approach is developed to reduce the complexity of the GD algorithm caused by parameter discretization. Additionally, the properties of the proposed algorithms are investigated, including convergence, complexity, and approximation error. The experimental results demonstrate that the performance of ERA is much better than that of the previous studies.

边缘智能模型拆分用户体验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。