提出边缘云协同推理框架,降低计算通信开销同时保持高性能。
ECCENTRIC: Edge-Cloud Collaboration Framework for Distributed Inference Using Knowledge Adaptation
- 通过边缘模型知识迁移优化云端模型,实现资源高效利用
- 在分类与目标检测任务中显著降低推理开销,性能接近最优
- 适合部署在资源受限的边缘设备与大规模云系统中
边缘AI的广泛应用推动了机器学习模型在多个领域的普及。尽管边缘系统具备较高的计算与通信效率,但受限于边缘设备的算力,多数情况下仍需依赖云端更强大的计算资源。然而,随着接入系统的边缘设备数量增加,云端推理带来的计算与通信成本急剧上升,形成计算、通信与性能之间的权衡。本文提出一种名为Eccentric的新框架,能够学习不同权衡程度的模型。该框架基于将边缘模型的知识迁移到云端模型的机制,在推理阶段有效降低系统计算与通信开销,同时尽可能维持最优性能。Eccentric可视为一种适用于边缘-云推理系统的新型压缩方法,能同步减少计算与通信负担。在分类和目标检测任务上的实证研究验证了该框架的有效性。
原文摘要 · Abstract (English)
The massive growth in the utilization of edge AI has made the applications of machine learning models ubiquitous in different domains. Despite the computation and communication efficiency of these systems, due to limited computation resources on edge devices, relying on more computationally rich systems on the cloud side is inevitable in most cases. Cloud inference systems can achieve the best performance while the computation and communication cost is dramatically increasing by the expansion of a number of edge devices relying on these systems. Hence, there is a trade-off between the computation, communication, and performance of these systems. In this paper, we propose a novel framework, dubbed as Eccentric that learns models with different levels of trade-offs between these conflicting objectives. This framework, based on an adaptation of knowledge from the edge model to the cloud one, reduces the computation and communication costs of the system during inference while achieving the best performance possible. The Eccentric framework can be considered as a new form of compression method suited for edge-cloud inference systems to reduce both computation and communication costs. Empirical studies on classification and object detection tasks corroborate the efficacy of this framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。