兼顾隐私与性能,优化边缘计算中模型部署与分配
Privacy-Aware Joint DNN Model Deployment and Partitioning Optimization for Collaborative Edge Inference Services
- 用李雅普诺夫方法将长期延迟最小化转化为每时隙可解问题
- 联合优化模型部署、用户-服务器关联和模型分割,降低推理延迟
- 适合关注边缘计算隐私与效率平衡的研究者或工程师
边缘推理(EI)作为应对云上深度神经网络(DNN)推理服务高响应延迟、扩展性差和严重数据隐私泄露的新兴范式,正受到广泛关注。然而,在资源受限的边缘设备上部署DNN模型带来了新挑战:计算/存储资源有限、服务需求动态变化、隐私风险加剧。本文提出一种新型隐私感知优化框架,联合优化DNN模型部署、用户-服务器关联与模型分割,以在资源与隐私约束下最小化长期平均推理延迟。该问题被建模为复杂且NP-hard的随机优化问题。为高效处理系统动态与计算复杂度,我们采用基于李雅普诺夫的方法,将长期目标转化为可处理的时隙级决策。此外,引入联盟形成博弈实现自适应用户-服务器关联,并设计贪心算法在每个联盟内完成模型部署。大量仿真表明,所提算法显著降低推理延迟,且始终满足隐私约束,在多种场景下均优于现有最优基线。
原文摘要 · Abstract (English)
Edge inference (EI) has emerged as a promising paradigm to address the growing limitations of cloud-based Deep Neural Network (DNN) inference services, such as high response latency, limited scalability, and severe data privacy exposure. However, deploying DNN models on resource-constrained edge devices introduces additional challenges, including limited computation/storage resources, dynamic service demands, and heightened privacy risks. To tackle these issues, this paper presents a novel privacy-aware optimization framework that jointly addresses DNN model deployment, user-server association, and model partitioning, with the goal of minimizing long-term average inference delay under resource and privacy constraints. The problem is formulated as a complex, NP-hard stochastic optimization. To efficiently handle system dynamics and computational complexity, we employ a Lyapunov-based approach to transform the long-term objective into tractable per-slot decisions. Furthermore, we introduce a coalition formation game to enable adaptive user-server association and design a greedy algorithm for model deployment within each coalition. Extensive simulations demonstrate that the proposed algorithm significantly reduces inference delay and consistently satisfies privacy constraints, outperforming state-of-the-art baselines across diverse scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。