让强化学习迁移更高效,结合策略与结构知识提升样本利用率
CADENT: Gated Hybrid Distillation for Sample-Efficient Transfer in Reinforcement Learning
- 融合策略指导与任务结构知识,动态加权教师建议
- 在稀疏奖励网格与连续控制任务中提升40%-60%样本效率
- 适合需要快速适应新环境的强化学习应用
迁移学习有望降低深度强化学习的高样本复杂性,但现有方法在源环境与目标环境间存在领域偏移时表现不佳。策略蒸馏提供有效的战术指导,却难以传递长期战略知识;基于自动机的方法可捕捉任务结构,但缺乏精细动作引导。本文提出上下文感知的体验门控迁移框架CADENT,将战略性的自动机知识与战术性的策略知识统一为连贯的指导信号。其核心创新在于经验门控的信任机制,可在状态-动作层面动态权衡教师指导与学生自身经验,实现对目标域特性的平滑适应。在从稀疏奖励网格世界到连续控制任务等挑战性环境中,CADENT相比基线方法实现40%-60%的样本效率提升,同时保持优异的最终性能,为强化学习中的自适应知识迁移提供了稳健方案。
原文摘要 · Abstract (English)
Transfer learning promises to reduce the high sample complexity of deep reinforcement learning (RL), yet existing methods struggle with domain shift between source and target environments. Policy distillation provides powerful tactical guidance but fails to transfer long-term strategic knowledge, while automaton-based methods capture task structure but lack fine-grained action guidance. This paper introduces Context-Aware Distillation with Experience-gated Transfer (CADENT), a framework that unifies strategic automaton-based knowledge with tactical policy-level knowledge into a coherent guidance signal. CADENT's key innovation is an experience-gated trust mechanism that dynamically weighs teacher guidance against the student's own experience at the state-action level, enabling graceful adaptation to target domain specifics. Across challenging environments, from sparse-reward grid worlds to continuous control tasks, CADENT achieves 40-60\% better sample efficiency than baselines while maintaining superior asymptotic performance, establishing a robust approach for adaptive knowledge transfer in RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。