通过自适应教师指导提升强化学习的安全性与迁移效率
Safety-Regulated Transfer Reinforcement Learning with Adaptive Teacher Guidance

- 根据安全成本动态调整教师干预,随学生安全性提升逐步放权
- 安全风险高时强化教师引导,达标后自动减弱,平均速度提升6.90%
- 基于策略兼容性重加权干预数据,稳定训练过程,适合安全敏感场景
我们提出安全调控自适应迁移强化学习(SRATRL),一种教师-学生框架,结合安全触发干预、安全自适应价值塑造及策略兼容性优化,实现高效目标域迁移。首先,设计安全触发闭环干预策略,根据即时安全成本激活教师指导,并依据学生近期安全表现自适应调整干预阈值,实现及时安全监督并逐步恢复学生自主性。其次,引入安全自适应教师引导的价值塑造方案,将教师一致性信号融入评判器目标,其贡献由安全约束乘子动态调节,在安全风险升高时增强引导,满足约束后逐步减弱。此外,提出教师-学生策略兼容性加权方法,按教师与学生策略对动作执行概率的相对差异重加权干预样本,缓解策略不匹配带来的负面优化影响。实验表明,相较于带有拉格朗日约束的近端策略优化基线,所提方法平均速度提升6.90%,碰撞率降低75.00%,在降低安全成本的同时保持良好任务效率。
原文摘要 · Abstract (English)
We propose Safety-Regulated Adaptive Transfer Reinforcement Learning (SRATRL), a teacher--student framework that combines safety-triggered intervention, safety-adaptive value shaping, and policy-compatibility-based optimization for efficient target-domain adaptation. First, a safety-triggered closed-loop intervention strategy is developed that activates teacher guidance according to the instantaneous safety cost and adaptively adjusts the intervention threshold based on the student policy's recent safety performance, thereby providing timely safety supervision while progressively restoring student autonomy as its safety improves. Next, a safety-adaptive teacher-guided value-shaping scheme is introduced, in which a teacher-consistency signal is incorporated into the critic target, and its contribution is dynamically regulated by the safety-constraint multiplier, enabling stronger teacher guidance under elevated safety risks and gradually weakening such guidance as the safety constraint is better satisfied. In addition, a teacher-student policy-compatibility weighting approach is proposed to alleviate the adverse optimization effects caused by policy mismatch. It reweights teacher-intervened transitions according to the relative likelihood of the executed action under the teacher and student policies, thereby improving policy-update stability. Experimental results demonstrate that compared with a Proximal Policy Optimization with Lagrangian constraint baseline, the proposed method improves the average velocity by 6.90%, and reduces the crash ratio by 75.00%. These results demonstrate that the proposed method can reduce safety costs while maintaining competitive task efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。