arXiv:2607.18288cs.LG2026-07

提出多时尺度强化学习框架,优化边缘云网络延迟与资源利用。

Multi-Timescale Latent-Action DRL for Joint Optimization in Edge-Cloud Networks

论文配图:Multi-Timescale Latent-Action DRL for Joint Optimization in Edge-Cloud Networks
图 1 · 摘自论文原文
  • 分长短期决策,用变分自编码压缩高维动作空间
  • 平均端到端延迟降低20.8%,资源利用率提升13%
  • 适合动态边缘计算场景的实时调度优化

在动态任务到达和异构资源条件下,边缘-云层级系统中的负载不均会显著增加排队延迟并降低资源利用率。为解决此问题,研究联合服务部署、计算委派与功率控制(JSCP)问题,以最小化平均端到端(e2e)延迟。该问题因离散与连续变量强耦合,属混合整数非凸且NP-hard。为实现可处理优化与稳定适应,基于决策动态差异,将问题分解为长期系统配置与短期资源分配子问题。在此基础上,提出基于潜在动作空间的双时尺度多层深度强化学习框架(2T-MDRL-LA),联合优化服务部署、用户关联、计算委派、任务卸载及用户发射功率。采用基于变分自编码器的潜在动作表示,高效压缩高维组合动作空间。仿真表明,所提框架能有效适应动态网络环境,性能接近分支定界解。相比无计算委派方案,平均e2e延迟降低20.8%,资源利用率提升13%,收敛速度比传统近端策略优化快约50%。

原文摘要 · Abstract (English)

Load imbalance across edge and cloud layers degrades latency performance in hierarchical edge-cloud computing (HECC) systems under dynamic task arrivals and heterogeneous resources, leading to severe queuing delays and inefficient resource utilization. To address this challenge, we study a joint service placement, computational delegation, and power control (JSCP) problem to minimize the average end-to-end (e2e) latency. The resulting JSCP problem is a mixed-integer nonconvex and NP-hard optimization problem due to the strong coupling between discrete and continuous variables. To enable tractable optimization and stable system adaptation, we exploit the inherent difference in decision dynamics and decompose the problem into long-term system configuration and short-term resource allocation subproblems. Based on this formulation, we propose a two-timescale multi-layer deep reinforcement learning framework with a latent action space (2T-MDRL-LA) to jointly optimize service placement, user association, computational delegation, task offloading, and user transmit power. A latent action representation based on a variational autoencoder is introduced to efficiently compress the high-dimensional combinatorial action space. Simulation results demonstrate that the proposed framework effectively adapts to dynamic network conditions and achieves near-optimal performance compared to branch-and-bound solutions. It achieves up to a 20.8% reduction in average e2e latency and a 13% improvement in resource utilization over the scheme without the computational delegation, while converging approximately 50% faster than conventional proximal policy optimization.

边缘计算强化学习资源优化时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。