arXiv:2506.09499cs.LGcs.AI2025-06

提出可组合可解释的策略核方程,让智能体高效规划长时序任务。

A Unified Theory of Compositionality, Modularity, and Interpretability in Markov Decision Processes

  • 用状态-时间策略核直接建模策略转移,替代传统价值函数
  • 支持跨任务复用、长期时空预测和语义事件概率记录
  • 适合需要可验证规划与内在动机的高维动态环境

我们提出了选项核贝尔曼方程(OKBEs),用于无奖励马尔可夫决策过程。不同于传统的值函数,OKBEs直接构建并优化一种称为状态-时间选项核(STOK)的预测映射,以最大化达成目标的概率同时避免违反约束。STOK是强化学习选项框架中具有可组合性、模块化和可解释性的策略转移核:1)可通过查普曼-科尔莫戈罗夫方程组合,实现多策略在长时域上的时空预测;2)高维STOK可高效以因子分解且可重构的形式表示与计算;3)记录可语义解释的目标达成与约束违反事件概率,支持形式化验证。针对难以求解的高维状态转移模型,可分解为局部STOK与目标条件策略,并聚合为因子分解的目标核,从而在目标层级上实现高维前向规划。这些特性使智能体具备高度灵活性,能快速合成元策略,跨任务复用规划表示,并通过赋能(empowerment)机制解释目标。我们认为,奖励最大化与可组合性、模块化及可解释性存在冲突;而OKBEs则促进这些性质,支持可验证的长时域规划与可扩展至动态高维世界模型的内在动机。

原文摘要 · Abstract (English)

We introduce Option Kernel Bellman Equations (OKBEs) for a new reward-free Markov Decision Process. Rather than a value function, OKBEs directly construct and optimize a predictive map called a state-time option kernel (STOK) to maximize the probability of completing a goal while avoiding constraint violations. STOKs are compositional, modular, and interpretable initiation-to-termination transition kernels for policies in the Options Framework of Reinforcement Learning. This means: 1) STOKs can be composed using Chapman-Kolmogorov equations to make spatiotemporal predictions for multiple policies over long horizons, 2) high-dimensional STOKs can be represented and computed efficiently in a factorized and reconfigurable form, and 3) STOKs record the probabilities of semantically interpretable goal-success and constraint-violation events, needed for formal verification. Given a high-dimensional state-transition model for an intractable planning problem, we can decompose it with local STOKs and goal-conditioned policies that are aggregated into a factorized goal kernel, making it possible to forward-plan at the level of goals in high-dimensions to solve the problem. These properties lead to highly flexible agents that can rapidly synthesize meta-policies, reuse planning representations across many tasks, and justify goals using empowerment, an intrinsic motivation function. We argue that reward-maximization is in conflict with the properties of compositionality, modularity, and interpretability. Alternatively, OKBEs facilitate these properties to support verifiable long-horizon planning and intrinsic motivation that scales to dynamic high-dimensional world-models.

强化学习可解释性规划选项框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。