arXiv:2502.14264cs.AI2025-02

用博弈论协调感知与决策,让智能体更高效地处理复杂视觉输入。

SPRIG: Stackelberg Perception-Reinforcement Learning with Internal Game Dynamics

  • 将感知与决策建模为合作型斯塔克尔伯格博弈,感知模块领导提取关键特征。
  • 在Atari BeamRider上比标准PPO高出约30%的回报,且有理论保障。
  • 适合研究感知-决策协同、强化学习架构设计的学者与工程师。

深度强化学习智能体常面临感知与决策组件难以有效协同的问题,尤其在高维感官输入环境中,特征重要性动态变化。本文提出SPRIG(基于内部博弈的斯塔克尔伯格感知-强化学习框架),将单个智能体内的感知-策略互动建模为合作型斯塔克尔伯格博弈:感知模块作为领导者,战略性地处理原始感官状态;策略模块作为跟随者,基于提取特征做出决策。SPRIG通过改进的贝尔曼算子提供理论保证,同时保留现代策略优化的优势。在Atari BeamRider环境上的实验表明,SPRIG相较标准PPO实现了约30%的回报提升,验证了其在特征提取与决策间实现博弈论平衡的有效性。

原文摘要 · Abstract (English)

Deep reinforcement learning agents often face challenges to effectively coordinate perception and decision-making components, particularly in environments with high-dimensional sensory inputs where feature relevance varies. This work introduces SPRIG (Stackelberg Perception-Reinforcement learning with Internal Game dynamics), a framework that models the internal perception-policy interaction within a single agent as a cooperative Stackelberg game. In SPRIG, the perception module acts as a leader, strategically processing raw sensory states, while the policy module follows, making decisions based on extracted features. SPRIG provides theoretical guarantees through a modified Bellman operator while preserving the benefits of modern policy optimization. Experimental results on the Atari BeamRider environment demonstrate SPRIG's effectiveness, achieving around 30% higher returns than standard PPO through its game-theoretical balance of feature extraction and decision-making.

强化学习感知协同博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。