前沿大推理模型在游戏学习中逼近人类行为与脑活动模式。
Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners

- 用人类游戏行为与脑影像数据联合评估模型,跨模态验证。
- 大推理模型预测脑活动比强化学习模型好一个数量级。
- 模型脑对齐依赖其对游戏状态的上下文表征,非后续推理。
人类在面对新环境时能快速习得抽象知识,并灵活运用以指导高效智能行为。现代人工智能系统能否实现类似能力?我们利用包含复杂人类游戏行为及同步fMRI记录的数据集,研究参与者学习需要规则发现、假设修正和多步规划的新视频游戏的过程。通过评估模型的游戏表现、人类学习行为匹配度以及对同一任务中脑活动的预测能力,对比了前沿大推理模型(LRMs)与无模型、有模型深度强化学习代理及贝叶斯理论代理。结果表明,前沿LRMs在游戏探索阶段最接近人类行为模式,且在皮层和皮层下区域的脑活动预测上优于强化学习方法一个数量级,且在置换控制下仍具稳健性。通过定向干预进一步发现,脑对齐反映的是模型对游戏状态的上下文表征,而非下游规划或推理过程。研究确立了大推理模型作为复杂自然环境中人类学习与决策的有力计算模型。项目页面含交互式回放:https://botcs.github.io/reason-to-play/
原文摘要 · Abstract (English)
Humans rapidly learn abstract knowledge when encountering novel environments and flexibly deploy this knowledge to guide efficient and intelligent action. Can modern AI systems learn and plan in a similar way? We study this question using a dataset of complex human gameplay with concurrent fMRI recordings, in which participants learn novel video games that require rule discovery, hypothesis revision, and multi-step planning. We jointly evaluate models by their ability to play the games, match human learning behavior, and predict brain activity during the same task, comparing a suite of frontier Large Reasoning Models (LRMs) against model-free and model-based deep reinforcement learning agents and a Bayesian theory-based agent. We find that frontier LRMs most closely match human behavioral patterns during game discovery and predict brain activity an order of magnitude better than both reinforcement learning alternatives across cortical and subcortical regions, with effects robust to permutation controls. Through targeted manipulations, we further show that brain alignment reflects the model's in-context representation of the game state rather than its downstream planning or reasoning. Our results establish LRMs as compelling computational accounts of human learning and decision making in complex, naturalistic environments. Project page with interactive replays: https://botcs.github.io/reason-to-play/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。