arXiv:2602.18797cs.DCcs.AI2026-02

用强化学习让物联网设备自主调度任务,降低碳排放并减少延迟。

Carbon-aware decentralized dynamic task offloading in MIMO-MEC networks via multi-agent reinforcement learning

  • 多个设备独立决策,基于本地信息动态选择任务是否本地处理或上传。
  • 碳排放最低且极端负载下丢包率接近零,优于传统方法。
  • 适合关注绿色计算和低延迟的物联网系统部署者。

海量物联网微服务需将可再生能源采集融入移动边缘计算(MEC),以构建可持续的eScience基础设施。任务到达的时空随机性与绿色能源供应的间歇性、多天线(MIMO)上行链路中的复杂用户间干扰,使得实时资源管理极为困难。传统集中式优化与离策略强化学习在密集网络中面临扩展性差与信令开销高的问题。本文提出基于多智能体近端策略优化的碳感知去中心化动态任务卸载框架CADDTO-PPO。将多用户MIMO-MEC系统建模为分布式部分可观测马尔可夫决策过程(DEC-POMDP),联合最小化碳排放、缓冲延迟及能量浪费。采用去中心化执行与参数共享(DEPS)架构,使物联网智能体仅依赖本地观测即可实现精细功率控制与卸载决策。此外,设计碳优先奖励机制,自适应优先选择绿色时段传输数据,实现系统吞吐量与电网碳足迹解耦。实验表明,CADDTO-PPO显著优于深度确定性策略梯度(DDPG)与李雅普诺夫基线。框架在极端流量负载下实现最低碳强度,并维持近零丢包率。架构分析验证其推理复杂度恒为O(1),具备未来可持续物联网部署的理论轻量化可行性。

原文摘要 · Abstract (English)

Massive internet of things microservices require integrating renewable energy harvesting into mobile edge computing (MEC) for sustainable eScience infrastructures. Spatiotemporal mismatches between stochastic task arrivals and intermittent green energy along with complex inter-user interference in multi-antenna (MIMO) uplinks complicate real-time resource management. Traditional centralized optimization and off-policy reinforcement learning struggle with scalability and signaling overhead in dense networks. This paper proposes CADDTO-PPO, a carbon-aware decentralized dynamic task offloading framework based on multi-agent proximal policy optimization. The multi-user MIMO-MEC system is modeled as a Decentralized Partially Observable Markov Decision Process (DEC-POMDP) to jointly minimize carbon emissions and buffer latency and energy wastage. A scalable architecture utilizes decentralized execution with parameter sharing (DEPS), which enables autonomous IoT agents to make fine-grained power control and offloading decisions based solely on local observations. Additionally, a carbon-first reward structure adaptively prioritizes green time slots for data transmission to decouple system throughput from grid-dependent carbon footprints. Finally, experimental results demonstrate CADDTO-PPO outperforms deep deterministic policy gradient (DDPG) and lyapunov-based baselines. The framework achieves the lowest carbon intensity and maintains near-zero packet overflow rates under extreme traffic loads. Architectural profiling validates the framework to demonstrate a constant $O(1)$ inference complexity and theoretical lightweight feasibility for future generation sustainable IoT deployments.

边缘计算强化学习低碳调度MIMO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。