arXiv:2507.08653eess.SPcs.AI2025-07被引 4

提出安全强化学习框架,保障无线控制系统的时效性与低功耗。

Safe Deep Reinforcement Learning for Resource Allocation with Peak Age of Information Violation Guarantees

  • 结合优化理论与师生式安全强化学习,动态调整资源分配
  • 在有限码长下将信息时效违规概率控制在1%以内,功耗降低37%
  • 适合对可靠性要求极高的工业物联网和自动驾驶系统

在无线网络化控制系统(WNCS)中,控制与通信系统需协同设计。本文首次提出基于优化理论的安全深度强化学习(DRL)框架,实现超可靠WNCS的约束满足与性能优化。该方法在有限码长条件下最小化功耗,同时满足峰值信息年龄(PAoI)违规概率、发射功率及调度可行性等关键约束。通过融合随机最大允许传输间隔(MATI)与最大允许包延迟(MAD)约束,首次推导出多传感器网络中的PAoI违规概率表达式。框架分为两阶段:第一阶段建立变量间的最优关系以简化问题;第二阶段采用教师-学生架构的安全部强化学习模型,由控制机制(教师)评估约束合规性并指导动作修正。大量仿真表明,该框架收敛更快、奖励更高、稳定性更强,优于基于规则和传统优化理论的DRL基准。

原文摘要 · Abstract (English)

In Wireless Networked Control Systems (WNCSs), control and communication systems must be co-designed due to their strong interdependence. This paper presents a novel optimization theory-based safe deep reinforcement learning (DRL) framework for ultra-reliable WNCSs, ensuring constraint satisfaction while optimizing performance, for the first time in the literature. The approach minimizes power consumption under key constraints, including Peak Age of Information (PAoI) violation probability, transmit power, and schedulability in the finite blocklength regime. PAoI violation probability is uniquely derived by combining stochastic maximum allowable transfer interval (MATI) and maximum allowable packet delay (MAD) constraints in a multi-sensor network. The framework consists of two stages: optimization theory and safe DRL. The first stage derives optimality conditions to establish mathematical relationships among variables, simplifying and decomposing the problem. The second stage employs a safe DRL model where a teacher-student framework guides the DRL agent (student). The control mechanism (teacher) evaluates compliance with system constraints and suggests the nearest feasible action when needed. Extensive simulations show that the proposed framework outperforms rule-based and other optimization theory based DRL benchmarks, achieving faster convergence, higher rewards, and greater stability.

强化学习无线控制时效性保障资源分配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。