arXiv:2605.27720cs.LGstat.AP2026-05

用贝叶斯方法评估强化学习着陆控制器的部署可行性。

Bayesian Deployment Approval for Learned Landing Controllers under Finite Rollout Validation

论文配图:Bayesian Deployment Approval for Learned Landing Controllers under Finite Rollout Validation
图 1 · 摘自论文原文
  • 基于贝叶斯后验推断量化策略在不确定条件下的着陆能力
  • 有限仿真下成功率达85%时仍可能误判,后验批准概率更可靠
  • 适合需要安全决策的自主系统部署评估,如无人机、航天器

强化学习与数据驱动的自主着陆控制器通常通过有限仿真轨迹中的累积奖励和经验成功率来评估。然而,这些经验指标未必能在不确定性下提供充分的部署依据。本文提出一种基于贝叶斯框架的部署审批方法,在有限回放证据下评估学习控制器的部署准备度。通过建立基于不确定工况下触地安全满足度的概率化着陆能力模型,并利用贝叶斯后验推断量化真实部署能力的不确定性。引入后验批准概率与后验部署风险,支持逐阶段验证中的批准/拒绝/继续决策。基于PPO与SAC控制器的仿真实验表明,仅优化成功次数和奖励可能导致在有限验证下的过度自信,而后验推理提供了更校准不确定性的部署评估。该框架实现了强化学习评估与不确定环境下部署导向验证之间的实用统计连接,可推广至更多类别的学习型自主系统。

原文摘要 · Abstract (English)

Reinforcement learning and data-driven autonomous controllers are commonly evaluated using cumulative reward and empirical success frequency under finite simulation trajectories. However, such empirical metrics do not necessarily provide sufficient statistical evidence regarding deployment readiness under uncertainty. This work develops a Bayesian approval framework for learned autonomous landing controllers under finite rollout evidence. A probabilistic landing capability formulation is introduced based on touchdown safety satisfaction under uncertain operating conditions, while Bayesian posterior inference is used to quantify uncertainty regarding the true deployment capability of learned policies. Posterior approval probability and posterior deployment risk are further introduced for deployment-oriented evaluation, together with a sequential validation framework supporting approve/reject/continue decisions during progressive rollout testing. Simulation experiments using PPO and SAC controllers demonstrate that empirical success and reward optimization may produce overconfident deployment interpretation under limited validation evidence, whereas posterior approval inference provides a more uncertainty-calibrated assessment of deployment readiness. The proposed framework provides a practical statistical connection between conventional reinforcement-learning evaluation and deployment-oriented validation under uncertainty and may be generalized to broader classes of learned autonomous systems.

强化学习部署评估贝叶斯方法自主系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。