arXiv:2603.08956econ.GNcs.LG2026-03综述被引 2

用强化学习解决高维经济问题,拓展动态规划的应用边界。

A Survey of Reinforcement Learning For Economics

  • 将强化学习作为动态规划的样本化延伸,处理高维状态与连续动作。
  • 在定价、库存、博弈等场景中验证算法有效性,但依赖精确仿真器。
  • 适合关注计算经济学、结构建模的学者,尤其对复杂决策问题感兴趣者。

本文重新向经济学家介绍强化学习方法。动态规划受维度诅咒限制,只能适用于小规模问题或可简化的大型问题;而越来越多的经济模型无法通过简化处理。强化学习算法为动态规划提供了自然的、基于样本的扩展,使其可应用于具有高维状态、连续动作和策略互动的复杂问题。本文梳理了经典规划与现代学习算法之间的理论联系,并通过定价、库存控制、战略博弈和偏好识别等模拟案例展示其机制。同时指出这些算法存在脆弱性、样本效率低、对超参数敏感以及缺乏全局收敛保证等问题,其成功受限于精确仿真器的存在。当结合经济结构时,强化学习提供了一个灵活但不完美的计算工具框架。附带的另一篇综述(Rust and Rawat, 2026b)探讨从观测行为推断偏好的逆问题。所有仿真代码公开可用。

原文摘要 · Abstract (English)

This survey (re)introduces reinforcement learning methods to economists. The curse of dimensionality limits how far exact dynamic programming can be effectively applied, forcing us to rely on suitably "small" problems or our ability to convert "big" problems into smaller ones. While this reduction has been sufficient for many classical applications, a growing class of economic models resists such reduction. Reinforcement learning algorithms offer a natural, sample-based extension of dynamic programming, extending tractability to problems with high-dimensional states, continuous actions, and strategic interactions. I review the theory connecting classical planning to modern learning algorithms and demonstrate their mechanics through simulated examples in pricing, inventory control, strategic games, and preference elicitation. I also examine the practical vulnerabilities of these algorithms, noting their brittleness, sample inefficiency, sensitivity to hyperparameters, and the absence of global convergence guarantees outside of tabular settings. The successes of reinforcement learning remain strictly bounded by these constraints, as well as a reliance on accurate simulators. When guided by economic structure, reinforcement learning provides a remarkably flexible framework. It stands as an imperfect, but promising, addition to the computational economist's toolkit. A companion survey (Rust and Rawat, 2026b) covers the inverse problem of inferring preferences from observed behavior. All simulation code is publicly available.

强化学习计算经济学动态规划模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。