提出可逆Simplex框架,实现目标导向强化学习的可信任务完成验证
Progress-Certified Reversible Simplex Supervision of Goal-Reaching Reinforcement Learning
- 结合冻结状态策略与目标级自适应恢复机制
- 实测20次任务全成功,债务检测违规率从93%降至0%
- 适合需要高可靠性的机器人控制与形式化验证场景
当状态聚合、模型失配和扰动导致标准强化学习转移无效时,任务完成难以验证。本文提出一种可逆Simplex框架,融合冻结有限状态策略、目标级鲁棒自适应恢复,并显式计算动作后的进度债务。向外圆化可达性验证受扰转移,界定被拒绝学习动作带来的债务,并构建鲁棒恢复与重入集合。若独立检查器同时接受两类证书,且恢复衰减超过验证债务,则每个完成的切换周期均减少存储,实现有限切换与有限样本目标进入。在受限端到端实例中,检查器解决所有保留义务,认证κ_L=0.482,确立ε_sw≥0.060。匹配数值实验显示:有债务门控时20/20达成目标,0/320违反债务测试;无则为18/20和742/796。24次硬件试验验证了1 kHz执行器堆栈上50毫秒监督性能;形式化认证仅限于所检模型。
原文摘要 · Abstract (English)
Task completion is difficult to certify when state aggregation, model mismatch, and disturbances invalidate nominal RL transitions. We present a reversible Simplex framework combining a frozen finite-state policy with goal-level robust-adaptive recovery and explicit accounting for post-action progress debt. Outward-rounded reachability verifies disturbed transitions, bounds debt from rejected learned actions, and constructs robust recovery and re-entry sets. If an independent checker accepts both certificates and the recovery decrement exceeds the verified debt, every completed switching cycle decreases storage, yielding finite switching and finite-sample goal entry. In a restricted end-to-end instance, the checker resolves all retained obligations, certifies $κ_L=0.482$, and establishes $\varepsilon_{\mathrm{sw}}\ge 0.060$. Matched numerical episodes yield 20/20 goal entries and 0/320 debt-test violations with debt gating, versus 18/20 and 742/796 without it. Twenty-four hardware trials evaluate 50 ms supervision above a 1 kHz actuator stack; formal certification remains limited to the checked model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。