从博弈到生成模型,探索强化学习在复杂决策中的统一机制
Reinforcement Learning: From Algorithms To Foundation Models

- 结合多智能体博弈与生成模型,构建目标驱动的决策框架
- 提出基于扩散模型的世界建模方法,支持高效视频生成与长期规划
- 适合关注智能体交互与基础模型融合的研究者
强化学习(RL)为在明确目标下的序列决策提供框架。经典形式中,RL研究智能体如何在动态环境中采取行动以最大化长期奖励。在更丰富的设定下,问题扩展至单个智能体与固定环境之外:智能行为可能需要策略互动、适应不确定性以及对高维世界的推理。本文从两个角度研究强化学习:游戏中的算法与基础模型时代的强化学习。第一部分聚焦于游戏中的多智能体强化学习,探讨激励、策略与均衡概念在竞争性与一般和环境中的交互,涵盖双人零和博弈、大规模视频游戏及具有通用结构的多智能体场景。这些工作研究多智能体系统中的学习及强化学习方法在交互环境中的表现。第二部分研究生成模型与基础模型下的强化学习,基于先验知识可丰富序列决策的设想。预训练生成模型与学习的世界模型作为表示工具和结构化先验,用于规划、控制与策略优化。本文开发了基于扩散的世界模型,研究高效视频生成的强化学习,探索生成模型作为策略类,以及动作塑造未来观测的交互式视频世界模型,并通过带记忆的架构实现长时序建模。这些贡献共同呈现了一个统一的强化学习视角:在复杂序列领域中以目标为导向的适应。从战略博弈到生成世界模型,本文强调强化学习如何连接决策、环境建模与新兴基础模型能力,为智能行为的基本原理提供了更广阔的视角。
原文摘要 · Abstract (English)
Reinforcement learning (RL) provides a framework for sequential decision making under explicit objectives. In its classical form, RL studies how an agent should act to maximise long-term reward in a dynamic environment. In richer settings, the problem extends beyond a single agent and fixed environment: intelligent behavior may require strategic interaction, adaptation to uncertainty, and reasoning over high-dimensional worlds. This thesis studies RL from two perspectives: algorithms in games and RL in the era of foundation models. The first part focuses on multi-agent RL in games. It examines how incentives, policies, and equilibrium concepts interact in competitive and general-sum environments, spanning two-player zero-sum games, large-scale video games, and multi-player settings with general structure. These works investigate learning in multi-agent systems and the behavior of RL methods in interactive environments. The second part studies RL with generative and foundation models, motivated by the idea that prior knowledge can enrich sequential decision making. Pretrained generative models and learned world models serve as representation tools and structured priors for planning, control, and policy optimization. The thesis develops diffusion-based world models, investigates RL for efficient video generation, explores generative models as policy classes, and studies interactive video world models in which actions shape future observations. It also addresses long-horizon modeling through architectures with memory. Together, these contributions present a unified view of RL as objective-driven adaptation in complex sequential domains. From strategic games to generative world models, the thesis highlights how RL connects decision making, environment modeling, and emerging foundation-model capabilities, offering a broader perspective on the principles underlying intelligent behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。