arXiv:2603.06565cs.AIcs.LG2026-03

用符号选项预训练提升深度强化学习的长期决策能力

Boosting deep Reinforcement Learning using pretraining with Logical Options

  • 通过符号选项预训练,引导神经策略避开短期奖励陷阱
  • 在长周期任务中表现优于主流神经、符号及混合方法
  • 适合需要目标导向行为的复杂强化学习场景

深度强化学习代理常因过早追求即时奖励而偏离目标。近期一些符号方法通过编码稀疏目标和对齐规划缓解此问题,但纯符号架构难扩展且不适用于连续环境。为此,我们提出一种受人类习得新技能启发的混合框架——混合层次化强化学习(H^2RL)。该方法在不牺牲深度策略表达力的前提下,注入符号结构:先通过逻辑选项进行预训练,引导策略摆脱短期奖励循环,转向目标驱动行为,再通过标准环境交互优化最终策略。实验表明,该方法在长周期决策任务中持续提升性能,显著优于强基准模型——包括神经、符号及神经符号方法。

原文摘要 · Abstract (English)

Deep reinforcement learning agents are often misaligned, as they over-exploit early reward signals. Recently, several symbolic approaches have addressed these challenges by encoding sparse objectives along with aligned plans. However, purely symbolic architectures are complex to scale and difficult to apply to continuous settings. Hence, we propose a hybrid approach, inspired by humans' ability to acquire new skills. We use a two-stage framework that injects symbolic structure into neural-based reinforcement learning agents without sacrificing the expressivity of deep policies. Our method, called Hybrid Hierarchical RL (H^2RL), introduces a logical option-based pretraining strategy to steer the learning policy away from short-term reward loops and toward goal-directed behavior while allowing the final policy to be refined via standard environment interaction. Empirically, we show that this approach consistently improves long-horizon decision-making and yields agents that outperform strong neural, symbolic, and neuro-symbolic baselines.

强化学习符号学习预训练目标导向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。