arXiv:2603.00129cs.MAcs.LG2026-03被引 1

提出安全协作框架,让边缘设备高效私密地协同完成深度模型推理。

Safe Multi-Agent Deep Reinforcement Learning for Privacy-Aware Edge-Device Collaborative DNN Inference

  • 分层强化学习框架,动态划分模型并分配资源
  • 满足严格延迟约束,同时降低能耗与隐私风险
  • 适合资源受限的隐私敏感型边缘计算场景

随着深度神经网络(DNN)推理在边缘和移动平台日益普及,隐私保护、资源限制和动态模型部署面临严峻挑战。本文提出一种隐私感知的协同推理框架,通过在边缘设备与服务器间自适应划分模型来应对。为在动态服务需求与资源约束下联合优化推理延迟、能耗与隐私成本,将问题建模为包含模型部署、用户-服务器关联、模型分割与资源分配的约束马尔可夫决策过程(CMDP)。提出分层约束多智能体近端策略优化算法(HC-MAPPO-L),基于拉格朗日松弛增强多智能体近端策略优化,实现长期延迟约束的自适应保障。通过三层结构分解CMDP:基于自回归的模型部署策略、带拉格朗日增强的用户关联与模型分割策略、注意力机制的资源分配策略。大量实验表明,该方法始终满足严格延迟约束,在能耗与隐私成本之间取得更优平衡,优于多种基准算法,且在不同规模与资源配置下表现稳健。

原文摘要 · Abstract (English)

As Deep Neural Network (DNN) inference becomes increasingly prevalent on edge and mobile platforms, critical challenges emerge in privacy protection, resource constraints, and dynamic model deployment. This paper proposes a privacy-aware collaborative inference framework, in which adaptive model partitioning is performed across edge devices and servers. To jointly optimize inference delay, energy consumption, and privacy cost under dynamic service demands and resource constraints, we formulate the joint problem as a Constrained Markov Decision Process (CMDP) that integrates model deployment, user-server association, model partitioning, and resource allocation. We propose a Hierarchical Constrained Multi-Agent Proximal Policy Optimization with Lagrangian relaxation (HC-MAPPO-L) algorithm, a safe reinforcement learning-based framework that enhances Multi-Agent Proximal Policy Optimization (MAPPO) with adaptive Lagrangian dual updates to enforce long-term delay constraints. To ensure tractability while maintaining coordination, we decompose the CMDP into three hierarchically structured policy layers: an auto-regressive based model deployment policy, a Lagrangian-enhanced user association and model partitioning policy, and an attention-based resource allocation policy. Extensive experimental results demonstrate that HC-MAPPO-L consistently satisfies stringent delay constraints while achieving a superior balance among energy consumption and privacy cost, outperforming representative baseline algorithms across varying problem scales and resource configurations.

边缘计算强化学习隐私保护协同推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。