arXiv:2608.17181cs.LGcs.GT2026-08

用势论视角重新理解强化学习,提升样本效率与理论约束。

Reinforcement Learning as (Discrete) Potential Theory

论文配图:Reinforcement Learning as (Discrete) Potential Theory
图 1 · 摘自论文原文
  • 将强化学习视为离散势论问题,统一建模状态价值
  • 在固定策略下实现更优的采样效率和理论可解释性
  • 框架可拓展至非线性场景,适合理论研究者

强化学习理论本质上依赖马尔可夫链的概率论基础。本文回顾概率论与势论之间的深层联系,探讨在固定策略假设下,核心强化学习表示与算法的势论视角。该视角可能为提升样本效率及施加形式化约束提供新路径。当放宽固定策略假设时,线性势论框架可自然推广至非线性情形。

原文摘要 · Abstract (English)

Reinforcement learning (RL) theory fundamentally depends on probability theory through the Markov chain. There is a deep connection between probability theory and potential theory. This paper reviews that connection and explores the potential-theoretic viewpoint for core reinforcement learning representations and algorithms under a fixed-policy assumption. This viewpoint may offer a path for improved sample efficiency and formal constraints that can be applied to RL. When the fixed-policy assumption is relaxed, the linear potential theory framework can be naturally extended to the nonlinear case.

强化学习势论理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。