系统性梳理延迟环境下强化学习控制方法,解决实际工程中的时延挑战。
Reinforcement Learning for Control Systems with Time Delays: A Comprehensive Survey
- 按延迟类型分类五类应对策略,涵盖状态扩展、记忆网络、预测模型等
- 指出不同方法在稳定性与适应性上的权衡,提供选型参考
- 适合研究网络化控制系统与多智能体协同的工程师和学者
过去十年中,强化学习(RL)在复杂动态系统的控制与决策中取得了显著进展。然而,大多数RL算法依赖马尔可夫决策过程假设,而实际的网络化物理系统常受感知延迟、执行滞后和通信约束影响,导致时延引入记忆效应,严重降低性能并破坏稳定性,尤其在多智能体环境中。本文系统综述了针对控制中时延问题的强化学习方法。首先形式化主要延迟类别,并分析其对马尔可夫性的破坏;随后将现有方法分为五大类:状态扩展与历史表征、带记忆的循环策略、基于预测与模型感知的方法、鲁棒及领域随机训练策略、具有显式约束处理的安全强化学习框架。对每类方法的原理、优势与局限进行剖析。通过对比分析揭示关键权衡,为不同延迟特征与安全要求下的方法选择提供指导。最后,提出未解挑战与未来方向,包括稳定性证明、大时延学习、多智能体通信协同设计与标准化基准测试。本综述旨在为延迟环境下开发可靠强化学习控制器的研究者与实践者提供统一参考。
原文摘要 · Abstract (English)
In the last decade, Reinforcement Learning (RL) has achieved remarkable success in the control and decision-making of complex dynamical systems. However, most RL algorithms rely on the Markov Decision Process assumption, which is violated in practical cyber-physical systems affected by sensing delays, actuation latencies, and communication constraints. Such time delays introduce memory effects that can significantly degrade performance and compromise stability, particularly in networked and multi-agent environments. This paper presents a comprehensive survey of RL methods designed to address time delays in control systems. We first formalize the main classes of delays and analyze their impact on the Markov property. We then systematically categorize existing approaches into five major families: state augmentation and history-based representations, recurrent policies with learned memory, predictor-based and model-aware methods, robust and domain-randomized training strategies, and safe RL frameworks with explicit constraint handling. For each family, we discuss underlying principles, practical advantages, and inherent limitations. A comparative analysis highlights key trade-offs among these approaches and provides practical guidelines for selecting suitable methods under different delay characteristics and safety requirements. Finally, we identify open challenges and promising research directions, including stability certification, large-delay learning, multi-agent communication co-design, and standardized benchmarking. This survey aims to serve as a unified reference for researchers and practitioners developing reliable RL-based controllers in delay-affected cyber-physical systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。