arXiv:2604.21464cs.LGcs.AI2026-04

用动态先验约束策略演化,让智能体决策更连贯稳定。

Dynamical Priors as a Training Objective in Reinforcement Learning

  • 引入外部状态动力学作为辅助损失,引导动作概率随时间平稳变化。
  • 在三个简单环境中,决策轨迹明显更结构化,避免突变与振荡。
  • 无需改环境或模型,即可控制智能体的时序行为特征,适合关注决策质量的研究者。

标准强化学习优化奖励但对决策随时间的演变约束极少,导致策略可能高分却表现出突发信心变化、振荡或停滞等不连贯行为。我们提出动态先验强化学习(DP-RL),在策略梯度学习中加入源自外部状态动力学的辅助损失,实现证据积累与滞后效应。不改变奖励、环境或策略结构,该先验可调控学习过程中动作概率的时间演化。在三个最小环境上,动态先验系统性地以任务相关方式改变决策轨迹,促进无法用通用平滑解释的时序结构化行为。结果表明,仅通过训练目标即可控制强化学习智能体决策的时序几何。

原文摘要 · Abstract (English)

Standard reinforcement learning (RL) optimizes policies for reward but imposes few constraints on how decisions evolve over time. As a result, policies may achieve high performance while exhibiting temporally incoherent behavior such as abrupt confidence shifts, oscillations, or degenerate inactivity. We introduce Dynamical Prior Reinforcement Learning (DP-RL), a training framework that augments policy gradient learning with an auxiliary loss derived from external state dynamics that implement evidence accumulation and hysteresis. Without modifying the reward, environment, or policy architecture, this prior shapes the temporal evolution of action probabilities during learning. Across three minimal environments, we show that dynamical priors systematically alter decision trajectories in task-dependent ways, promoting temporally structured behavior that cannot be explained by generic smoothing. These results demonstrate that training objectives alone can control the temporal geometry of decision-making in RL agents.

强化学习决策时序动态先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。