arXiv:2509.01432cs.LG2025-09

用几何视角统一强化学习中的目标优化问题。

The Geometry of Nonlinear Reinforcement Learning

  • 将多种目标视为行为空间上的统一优化问题。
  • 自然推广经典算法至非线性效用与凸约束场景。
  • 适合研究深度强化学习与几何结合的学者。

奖励最大化、安全探索和内在动机在强化学习中常被当作独立目标研究。本文提出一个统一的几何框架,将这些目标视为环境可达长期行为空间上的单一优化问题。在此框架下,策略镜像下降、自然策略梯度和信任区域算法等经典方法可自然推广至非线性效用和凸约束情形。该视角有效捕捉了鲁棒性、安全性、探索性和多样性等目标,并指出了几何与深度强化学习交叉领域的开放挑战。

原文摘要 · Abstract (English)

Reward maximization, safe exploration, and intrinsic motivation are often studied as separate objectives in reinforcement learning (RL). We present a unified geometric framework, that views these goals as instances of a single optimization problem on the space of achievable long-term behavior in an environment. Within this framework, classical methods such as policy mirror descent, natural policy gradient, and trust-region algorithms naturally generalize to nonlinear utilities and convex constraints. We illustrate how this perspective captures robustness, safety, exploration, and diversity objectives, and outline open challenges at the interface of geometry and deep RL.

强化学习几何优化非线性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。