arXiv:2605.24759cs.LG2026-05

用带约束的反馈环重定义强化学习,让算法更稳定可靠。

A Contractive Feedback Semantics for Reinforcement Learning

  • 将决策过程视为可组合的开放组件,通过收缩反馈环实现无限时域评估。
  • 证明局部误差在特定电路中受控,状态抽象保持值不变并给出误差上限。
  • 适合研究强化学习形式化、安全控制与理论验证的学者参考。

折扣强化学习通常通过闭合马尔可夫决策过程上的贝尔曼方程来表述。本文提出一种组合视角:将单步决策过程视为开放的随机组件,通过闭合一个收缩的反馈环来获得无限时域策略评估。由此产生的语义为开放组件分配类型化的贝尔曼变换器,将串行与并行连接解释为变换器的组合与张量积,将反馈解释为由唯一不动点实现的可接受有界巴拿赫迹。该视角带来三个理论结果:第一,在允许的良类型有界单孔上下文中,近似组件等价是上下文同余,局部算子误差在反馈节点满足一致有界性条件下仍受控;第二,精确与近似状态抽象成为可交换或近似交换的余代数图,给出价值保持性和显式的上确界范数畸变界;第三,在单调ω-连续收缩变换器语义下,安全、风险与资源规范可表示为量化值契约,局部归纳界通过连线与反馈由最小不动点推理传递。核心主张并非所有强化学习态射构成全局迹幺半群范畴,而是折扣贝尔曼评估在允许的有界电路类上具有收缩反馈语义。

原文摘要 · Abstract (English)

Discounted reinforcement learning is usually presented through Bellman equations on closed Markov decision processes. This paper develops a compositional view: a one-step decision process is treated as an open stochastic component, and infinite-horizon policy evaluation is obtained by closing a contractive feedback loop. The resulting semantics assigns typed Bellman transformers to open components, interprets series and parallel wiring as composition and tensoring of transformers, and interprets feedback as an admissible guarded Banach trace realized by a unique fixed point. This perspective yields three theoretical consequences. First, approximate component equivalence is a contextual congruence for admitted well-typed guarded one-hole contexts: local operator error remains controlled after plugging the component into a surrounding circuit that uses the hole once and whose feedback nodes have certified uniform guardedness. Second, exact and approximate state abstractions become commuting or near-commuting coalgebraic diagrams, giving value-preservation and explicit sup-norm distortion bounds. Third, under monotone $ω$-continuous contract-transformer semantics, safety, risk, and resource specifications can be represented as quantale-valued contracts, where local inductive bounds lift through wiring and feedback by least-fixed-point reasoning. Its central claim is not that all RL morphisms form a global traced monoidal category, but that discounted Bellman evaluation admits a contractive feedback semantics on the admissible class of guarded circuits.

强化学习形式语义反馈系统稳定性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。