arXiv:2607.16210cs.AI2026-07中稿 · the 35th Internati…综述

梳理强化学习策略验证方法,帮开发者判断模型是否安全可靠。

A Survey on the Verification of Reinforcement Learning Policies

论文配图:A Survey on the Verification of Reinforcement Learning Policies
图 1 · 摘自论文原文
  • 按形式化/概率性、单步/多步、保证强度三维度分类验证方法
  • 揭示现有方法的理论基础与隐含假设,统一分析框架
  • 适合关注RL安全性的研究人员和工业部署者参考

强化学习在复杂安全关键领域应用日益广泛,但基于神经网络的策略缺乏严格的行为保障,已成为部署的主要障碍。随着策略表达能力和规模的提升,这一挑战愈发突出,相关研究迅速增长但概念碎片化严重。本文提供一个统一视角,构建涵盖验证范式(形式化与概率性)、时间范围(单步与多步)和保证强度三个维度的分类体系,统合底层理论基础,明确隐含假设与局限性,并指出未来发展方向。

原文摘要 · Abstract (English)

Reinforcement learning (RL) is increasingly applied in complex, safety-critical domains, yet the lack of rigorous behavioral guarantees for neural network-based policies remains a major barrier to deployment. Recent advances in policy expressiveness and scale have intensified this challenge, leading to a rapidly growing but conceptually fragmented body of work on RL policy verification. This survey provides a unifying perspective on RL verification methods. We introduce a taxonomy that clarifies relationships among existing approaches along three axes: verification paradigm (formal versus probabilistic), temporal scope (step-wise versus multi-step), and guarantees strength. Beyond taxonomy, we unify underlying theoretical foundations, make implicit assumptions and limitations explicit, and identify emerging directions.

强化学习策略验证安全保证综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。