为深度演员-评论家算法提供可实用的风险证书,仅需少量测试数据。
Deep Actor-Critics with Tight Risk Certificates
- 用递归PAC-Bayes理论,结合少量预训练策略的测试轨迹生成风险证书。
- 在多个运动任务中,证书准确预测泛化性能,误差极小。
- 适合关注强化学习部署安全性的研究者与工程师。
深度演员-评论家算法已深刻影响日常应用,尤其推动大语言模型通过用户反馈持续优化。然而其在物理系统中的部署仍受限,因缺乏对故障风险的量化验证机制。本文证明,可通过验证期观测数据,为深度演员-评论家算法构建紧致的风险证书,以预测泛化性能。核心洞察在于:极少量可行的评估轨迹(来自预训练策略)即可生成高精度证书。我们采用一种新型递归PAC-Bayes方法,将验证数据分段,递归构建每段预测器的超额损失上界,并以前一段预测器作为数据驱动的先验。在多个运动任务、多种演员-评论家方法及不同策略成熟度下,实验表明所生成的风险证书足够紧致,具备实际部署潜力。
原文摘要 · Abstract (English)
Deep actor-critic algorithms have reached a level where they influence everyday life. They are a driving force behind continual improvement of large language models through user feedback. However, their deployment in physical systems is not yet widely adopted, mainly because no validation scheme fully quantifies their risk of malfunction. We demonstrate that it is possible to develop tight risk certificates for deep actor-critic algorithms that predict generalization performance from validation-time observations. Our key insight centers on the effectiveness of minimal evaluation data. A small feasible set of evaluation roll-outs collected from a pretrained policy suffices to produce accurate risk certificates when combined with a simple adaptation of PAC-Bayes theory. Specifically, we adopt a recently introduced recursive PAC-Bayes approach, which splits validation data into portions and recursively builds PAC-Bayes bounds on the excess loss of each portion's predictor, using the predictor from the previous portion as a data-informed prior. Our empirical results across multiple locomotion tasks, actor-critic methods, and policy expertise levels demonstrate risk certificates tight enough to be considered for practical use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。