用数学证明提前防风险,让无线系统又安全又自主。
A Safety-Constrained Reinforcement Learning Framework for Reliable Wireless Autonomy

- 用轻量级数学证书实时验证动作是否安全
- 通过授权预算控制干预频率,平衡安全与效率
- 在无人机、车联网等高可靠场景特别有用
人工智能与强化学习在无线系统中展现出巨大潜力,可实现动态频谱分配、流量管理及大规模物联网协同。然而,在关键任务应用中部署时可能引发不安全行为,如无人机碰撞、拒绝服务事件或车载网络不稳定。现有安全机制多为被动响应,依赖异常检测或回滚控制器,仅在危险行为发生后介入,无法保障超可靠低时延通信(URLLC)环境下的可靠性。本文提出一种主动式安全约束强化学习框架,结合证明携带控制(PCC)与赋能预算(EB)执行机制。每个智能体动作均通过轻量级数学证书验证,确保符合干扰约束;赋能预算则调控安全干预频率,平衡安全性与自主性。我们在基于近端策略优化(PPO)的上行链路调度任务中实现了该框架。仿真结果表明,所提PCC+EB控制器完全消除不安全传输,同时保持系统吞吐量和可预测自主性。相比无约束及被动基线方法,本方法在极小性能损失下实现可证明的安全保障,凸显其在6G未来无线自治系统中的可信应用潜力。
原文摘要 · Abstract (English)
Artificial intelligence (AI) and reinforcement learning (RL) have shown significant promise in wireless systems, enabling dynamic spectrum allocation, traffic management, and large-scale Internet of Things (IoT) coordination. However, their deployment in mission-critical applications introduces the risk of unsafe emergent behaviors, such as UAV collisions, denial-of-service events, or instability in vehicular networks. Existing safety mechanisms are predominantly reactive, relying on anomaly detection or fallback controllers that intervene only after unsafe actions occur, which cannot guarantee reliability in ultra-reliable low-latency communication (URLLC) settings. In this work, we propose a proactive safety-constrained RL framework that integrates proof-carrying control (PCC) with empowerment-budgeted (EB) enforcement. Each agent action is verified through lightweight mathematical certificates to ensure compliance with interference constraints, while empowerment budgets regulate the frequency of safety overrides to balance safety and autonomy. We implement this framework on a wireless uplink scheduling task using Proximal Policy Optimization (PPO). Simulation results demonstrate that the proposed PCC+EB controller eliminates unsafe transmissions while preserving system throughput and predictable autonomy. Compared with unconstrained and reactive baselines, our method achieves provable safety guarantees with minimal performance degradation. These results highlight the potential of proactive safety constrained RL to enable trustworthy wireless autonomy in future 6G networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。