跨边云部署深度学习模型,优化延迟、隐私与成本的权衡。
Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
- 从多目标优化视角整合边云资源,动态分配模型推理任务。
- 在保证精度前提下,降低延迟与通信开销,提升隐私保护能力。
- 适合研究边缘计算与AI部署的学者及工程师参考。
随着物联网和移动设备的快速发展,虚拟现实、增强现实及基于语言模型的聊天机器人等边缘智能应用日益普及。然而,受限于边缘设备的计算能力,难以支撑日益庞大复杂的深度学习模型。为此,研究者提出将深度学习模型分片优化并跨用户设备、边缘服务器与云端进行卸载。在此架构下,用户可利用不同层级的服务:边缘资源提供低响应延迟,云端则以较低成本支持计算密集型任务。但模型分片间的通信可能引发传输瓶颈,并带来数据泄露风险。近期研究致力于在模型精度、计算延迟、传输延迟与隐私保护之间取得平衡,采用模型压缩、模型蒸馏、传输压缩及模型结构适配(如内部分类器)等技术。本文综述了当前主流的模型卸载方法与适应性技术,系统分析其对推断延迟、数据隐私与资源成本构成的多目标优化问题的影响。
原文摘要 · Abstract (English)
Edge intelligent applications like VR/AR and language model based chatbots have become widespread with the rapid expansion of IoT and mobile devices. However, constrained edge devices often cannot serve the increasingly large and complex deep learning (DL) models. To mitigate these challenges, researchers have proposed optimizing and offloading partitions of DL models among user devices, edge servers, and the cloud. In this setting, users can take advantage of different services to support their intelligent applications. For example, edge resources offer low response latency. In contrast, cloud platforms provide low monetary cost computation resources for computation-intensive workloads. However, communication between DL model partitions can introduce transmission bottlenecks and pose risks of data leakage. Recent research aims to balance accuracy, computation delay, transmission delay, and privacy concerns. They address these issues with model compression, model distillation, transmission compression, and model architecture adaptations, including internal classifiers. This survey contextualizes the state-of-the-art model offloading methods and model adaptation techniques by studying their implication to a multi-objective optimization comprising inference latency, data privacy, and resource monetary cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。