用大模型提升强化学习的适应性与可解释性
LLM-assisted Semantic Option Discovery for Facilitating Adaptive Deep Reinforcement Learning
- 将自然语言指令转为可执行规则,实现语义化技能复用
- 在两个环境中数据效率提升,约束合规率显著提高
- 适合需要安全可控、可解释强化学习的应用场景
尽管深度强化学习在复杂任务中取得显著进展,但在实际应用中仍面临数据效率低、可解释性差和跨环境迁移能力弱等关键问题。现有策略对环境变化敏感,难以保证行为安全与合规。近期研究显示,结合大语言模型与符号规划具有潜力。受此启发,我们提出一种新型大模型驱动的闭环框架,通过将自然语言指令映射为可执行规则,并对自动生成的选项进行语义标注,实现语义驱动的技能复用与实时约束监控。该方法利用大模型的通用知识提升探索效率,支持相似环境下的可迁移选项,并通过语义注释提供内在可解释性。我们在 Office World 与 Montezuma's Revenge 两个领域进行实验,结果表明其在数据效率、约束合规性和跨任务迁移能力方面均表现更优。
原文摘要 · Abstract (English)
Despite achieving remarkable success in complex tasks, Deep Reinforcement Learning (DRL) is still suffering from critical issues in practical applications, such as low data efficiency, lack of interpretability, and limited cross-environment transferability. However, the learned policy generating actions based on states are sensitive to the environmental changes, struggling to guarantee behavioral safety and compliance. Recent research shows that integrating Large Language Models (LLMs) with symbolic planning is promising in addressing these challenges. Inspired by this, we introduce a novel LLM-driven closed-loop framework, which enables semantic-driven skill reuse and real-time constraint monitoring by mapping natural language instructions into executable rules and semantically annotating automatically created options. The proposed approach utilizes the general knowledge of LLMs to facilitate exploration efficiency and adapt to transferable options for similar environments, and provides inherent interpretability through semantic annotations. To validate the effectiveness of this framework, we conduct experiments on two domains, Office World and Montezuma's Revenge, respectively. The results demonstrate superior performance in data efficiency, constraint compliance, and cross-task transferability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。