arXiv:2607.13160cs.ITcs.AI2026-07

用可重构智能表面优化边缘计算能效与延迟,实现快速决策。

Active Beyond-Diagonal RIS Empowered Heterogeneous Edge Computing: A Distributional Reinforcement Learning Approach

论文配图:Active Beyond-Diagonal RIS Empowered Heterogeneous Edge Computing: A Distributional Reinforcement Learning Approach
图 1 · 摘自论文原文
  • 基于分布强化学习,联合优化资源分配与信号反射策略。
  • 在81.67%可行性下达成最优能-时平衡,单次决策仅需0.0267秒。
  • 适合复杂动态环境下的移动边缘计算系统部署。

主动超对角可重构智能表面(BD-RIS)通过收发混合模式实现信号增强与全向覆盖,为异构移动边缘计算(MEC)系统的抗遮挡上行卸载提供可行方案。然而,实际中由互易器件实现的混合模式会引入跨扇区能量泄漏,重塑系统级能-时权衡。本文研究互易主动BD-RIS辅助下异构MEC的能耗感知卸载与资源分配,其中卸载决策、CPU/GPU计算分配、发射功率、接收处理及主动BD-RIS配置紧密耦合。该问题为高维混合整数非凸优化,传统逐实例优化难以高效求解。为此,提出基于改进分布软演员-评论家算法(DSAC-T)的端到端联合优化框架。通过建模回报分布而非仅期望值,提升在奖励异质性与可行性边界敏感情况下的策略稳定性。相比其他基线算法,DSAC-T在81.67%可行性率下取得最佳能-时奖励,且单场景在线决策时间仅为0.0267秒。

原文摘要 · Abstract (English)

Active beyond-diagonal reconfigurable intelligent surfaces (BD-RISs) enables hybrid transmitting and reflecting mode to achieve effective signal amplification and full-space coverage, thus providing a promising solution for blockage-aware uplink offloading in heterogeneous mobile edge computing (MEC) systems. However, practical hybrid mode active BD-RIS are realized by reciprocal devices, which inherently generate cross-sector energy leakage that will reshape the system-level energy-latency tradeoff. This paper studies energy-aware offloading and resource allocation for reciprocal active BD-RIS-assisted heterogeneous MEC, where offloading decisions, CPU/GPU computation allocation, transmit powers, receive processing, and active BD-RIS are tightly coupled. The resulting problem is a high-dimensional mixed integer nonconvex problem and is difficult to solve efficiently by conventional per-instance optimization. To address this challenge, we develop an end-to-end joint optimization framework based on a refined version of the distributional soft actor--critic algorithm, named as DSAC-T. By modeling return distributions rather than only expected values, DSAC-T improves policy stability under reward heterogeneity and feasibility-boundary sensitivity. Compared with other baseline algorithms, DSAC-T achieves the best energy-latency reward, the highest feasibility ratio of 81.67%, and a fast online decision time of 0.0267 s per scenario.

边缘计算强化学习智能表面资源分配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。