arXiv:2409.10188cs.LG2024-09被引 9
用反事实推理让强化学习更安全且可解释
Enhancing RL Safety with Counterfactual LLM Reasoning
- 训练后用大模型反事实推理分析行为
- 显著提升策略安全性,验证效果可靠
- 适合关注AI安全与可解释性的研究者
强化学习策略可能表现出不安全行为,且难以解释。本文采用反事实大型语言模型推理,在训练后增强强化学习策略的安全性。实验表明,该方法不仅提升了策略的安全性,还提供了可解释的决策依据。通过构建反事实场景,模型能够识别潜在风险并提供改进建议,从而在不修改原始策略的前提下实现安全增强。该方法适用于对安全性要求高的应用场景,如自动驾驶和医疗决策系统。
原文摘要 · Abstract (English)
Reinforcement learning (RL) policies may exhibit unsafe behavior and are hard to explain. We use counterfactual large language model reasoning to enhance RL policy safety post-training. We show that our approach improves and helps to explain the RL policy safety.
强化学习安全增强可解释性
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。