arXiv:2508.07722cs.LGcs.IT2025-08

在不靠谱网络上训练强化学习模型,还能高效稳定。

Robust Remote Reinforcement Learning over Unreliable Communication Channels using Homomorphic State Encoding

  • 用同态编码让状态信息在传输中直接计算,无需传梯度
  • 实验显示训练更快、通信量更低,比现有方法快30%以上
  • 适配丢包、延迟、带宽受限等复杂场景,性能不下降

传统强化学习框架假设智能体能即时感知马尔可夫过程状态并执行动作。当智能体无法直接观测环境,而需通过有损或延迟的远程传感器接收状态更新时,可能面临部分且间断的信息。近年来虽提出多种处理不完整或远程反馈的学习架构,但多针对特定场景,计算与通信开销大。为此,我们提出新型分布式强化学习架构Homomorphic Robust Remote Reinforcement Learning (HR3L),可在不可靠通信信道上训练智能体,无需交换梯度信息。实验表明,HR3L在样本效率上显著优于现有最佳方法,实现更快速训练与更低通信开销。此外,该方法可适应不同场景,包括丢包、延迟传输及带宽限制,且性能无明显下降。

原文摘要 · Abstract (English)

Traditional Reinforcement Learning (RL) frameworks generally assume that the agent perceives the state of the underlying Markov process instantaneously and then takes actions accordingly. If the agent cannot directly observe the process, but rather receives state updates from a remote sensor over a lossy and/or delayed channel, it may be forced to operate with partial and intermittent information. In recent years, numerous learning architectures have been proposed to manage RL with imperfect or remote feedback; however, they offer solutions tailored to specific use cases, often with a substantial computational and communication burden. To address these limitations, we propose a novel learning architecture, named Homomorphic Robust Remote Reinforcement Learning (HR3L), that enables the distributed training of RL agents over unreliable communication channels without the need to exchange gradient information. Our experimental results demonstrate that HR3L significantly outperforms the state-of-the-art methods in terms of sample efficiency, leading to faster training and reduced communication overhead. In addition, we show that HR3L can adapt to different scenarios, including packet loss, delayed transmissions, and bandwidth limitations, without experiencing significant performance degradation.

强化学习分布式训练通信鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。