用逻辑约束指导强化学习,实现大规模系统安全高效优化
Logic-informed reinforcement learning for cross-domain optimization of large-scale cyber-physical systems
- 将一阶逻辑动态定义可行动作流形,直接投影生成合法动作
- 工业装配系统中能耗与工期联合目标降低36.47%~44.33%,零违规
- 逻辑表达式可跨领域迁移,适合工业制造、交通信号等场景
网络物理系统(CPS)需在严格的安全逻辑约束下联合优化离散的网络动作与连续的物理参数。现有分层方法常牺牲全局最优性,而混合动作空间的强化学习依赖脆弱的奖励惩罚、掩码或屏蔽机制,难以保证约束满足。本文提出逻辑引导强化学习(LIRL),为标准策略梯度算法引入投影机制,将低维隐式动作映射到由一阶逻辑即时定义的可行动作流形上,确保每一步探索均合法且无需调参。在工业制造、电动汽车充电站及交通信号控制等多个场景中验证,所提方法优于现有分层优化方案。以工业机器人减速器装配为例,相较于传统分层调度方法,联合完工时间-能耗目标最多降低36.47%至44.33%,始终零约束违反,并显著超越当前先进混合动作强化学习基线。由于采用声明式逻辑约束形式,该框架可无缝迁移到智能交通、智能电网等领域,为大规模CPS的安全实时优化铺平道路。
原文摘要 · Abstract (English)
Cyber-physical systems (CPS) require the joint optimization of discrete cyber actions and continuous physical parameters under stringent safety logic constraints. However, existing hierarchical approaches often compromise global optimality, whereas reinforcement learning (RL) in hybrid action spaces often relies on brittle reward penalties, masking, or shielding and struggles to guarantee constraint satisfaction. We present logic-informed reinforcement learning (LIRL), which equips standard policy-gradient algorithms with projection that maps a low-dimensional latent action onto the admissible hybrid manifold defined on-the-fly by first-order logic. This guarantees feasibility of every exploratory step without penalty tuning. Experimental evaluations have been conducted across multiple scenarios, including industrial manufacturing, electric vehicle charging stations, and traffic signal control, in all of which the proposed method outperforms existing hierarchical optimization approaches. Taking a robotic reducer assembly system in industrial manufacturing as an example, LIRL achieves a 36.47\% to 44.33\% reduction at most in the combined makespan-energy objective compared to conventional industrial hierarchical scheduling methods. Meanwhile, it consistently maintains zero constraint violations and significantly surpasses state-of-the-art hybrid-action reinforcement learning baselines. Thanks to its declarative logic-based constraint formulation, the framework can be seamlessly transferred to other domains such as smart transportation and smart grid, thereby paving the way for safe and real-time optimization in large-scale CPS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。