用强化学习模拟经济行为时,如何让智能体从操纵者变竞争者。
From Individual Learning to Market Equilibrium: Correcting Structural and Parametric Biases in RL Simulations of Economic Models
- 设计均值场强化学习框架,让智能体在固定宏观经济场中学习
- 修正了经济折现与强化学习之间的参数偏差,使策略收敛至均衡
- 适合研究学习型经济主体的计算社会科学研究者
强化学习应用于经济建模时,暴露出均衡理论假设与学习智能体涌现行为的根本矛盾。传统经济模型假设个体是市场条件的接受者,而简单的单智能体强化学习则激励其成为环境操纵者。本文在凹生产搜索匹配模型中首次揭示此差异:标准强化学习智能体学得非均衡、买方垄断策略。此外,我们识别出经济折现与强化学习对跨期成本处理不一致导致的参数偏差。为此,提出校准的均值场强化学习框架,将代表性智能体置于固定宏观经济场中,并调整代价函数以反映经济机会成本。迭代算法收敛至自洽不动点,使智能体策略与竞争均衡一致。该方法为计算社会科学中建模学习型经济主体提供了可操作且理论严谨的路径。
原文摘要 · Abstract (English)
The application of Reinforcement Learning (RL) to economic modeling reveals a fundamental conflict between the assumptions of equilibrium theory and the emergent behavior of learning agents. While canonical economic models assume atomistic agents act as `takers' of aggregate market conditions, a naive single-agent RL simulation incentivizes the agent to become a `manipulator' of its environment. This paper first demonstrates this discrepancy within a search-and-matching model with concave production, showing that a standard RL agent learns a non-equilibrium, monopsonistic policy. Additionally, we identify a parametric bias arising from the mismatch between economic discounting and RL's treatment of intertemporal costs. To address both issues, we propose a calibrated Mean-Field Reinforcement Learning framework that embeds a representative agent in a fixed macroeconomic field and adjusts the cost function to reflect economic opportunity costs. Our iterative algorithm converges to a self-consistent fixed point where the agent's policy aligns with the competitive equilibrium. This approach provides a tractable and theoretically sound methodology for modeling learning agents in economic systems within the broader domain of computational social science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。