用可解释强化学习优化桥梁全生命周期维护决策
Interpretable Deep Reinforcement Learning for Element-level Bridge Life-cycle Optimization
- 采用可微分的斜决策树作为策略函数近似器
- 生成节点少、深度浅的可读决策树,接近最优策略
- 适合需要透明化决策的桥梁管理机构使用
自2022年起实施的《国家桥梁普查新规范》(SNBI)强调基于构件级状态(CS)进行风险导向的桥梁管理。与传统整体部件评分不同,构件级状态数据通过四个维度的概率数组(即状态比例)表示桥梁状况。这种细粒度数据虽提升表征能力,却因状态空间从单一整数扩展为四维概率分布,给制定最优全生命周期维护策略带来挑战。本文提出一种新型可解释强化学习方法,基于构件级状态表示寻求最优维护策略。相比现有方法,该算法生成的策略以确定性斜决策树形式呈现,节点数量和深度合理,便于人类理解与审计,并可直接集成至现有桥梁管理系统。为实现近最优策略,提出三项关键改进:(a) 使用可微分软决策树作为策略网络近似器;(b) 训练中引入温度退火机制;(c) 结合正则化与剪枝规则控制策略复杂度。三者协同可生成高可解释性策略。在监督学习与强化学习设置下验证了各技术的优劣权衡。框架在钢梁桥全生命周期优化问题中得到应用演示。
原文摘要 · Abstract (English)
The new Specifications for the National Bridge Inventory (SNBI), in effect from 2022, emphasize the use of element-level condition states (CS) for risk-based bridge management. Instead of a general component rating, element-level condition data use an array of relative CS quantities (i.e., CS proportions) to represent the condition of a bridge. Although this greatly increases the granularity of bridge condition data, it introduces challenges to set up optimal life-cycle policies due to the expanded state space from one single categorical integer to four-dimensional probability arrays. This study proposes a new interpretable reinforcement learning (RL) approach to seek optimal life-cycle policies based on element-level state representations. Compared to existing RL methods, the proposed algorithm yields life-cycle policies in the form of oblique decision trees with reasonable amounts of nodes and depth, making them directly understandable and auditable by humans and easily implementable into current bridge management systems. To achieve near-optimal policies, the proposed approach introduces three major improvements to existing RL methods: (a) the use of differentiable soft tree models as actor function approximators, (b) a temperature annealing process during training, and (c) regularization paired with pruning rules to limit policy complexity. Collectively, these improvements can yield interpretable life-cycle policies in the form of deterministic oblique decision trees. The benefits and trade-offs from these techniques are demonstrated in both supervised and reinforcement learning settings. The resulting framework is illustrated in a life-cycle optimization problem for steel girder bridges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。