arXiv:2412.03111cs.AI2024-12被引 2

用元认知强化学习模拟人类如何发现新规划策略。

Experience-driven discovery of planning strategies

  • 通过元认知强化学习机制探索新规划策略
  • 模型能解释人类策略发现行为,但速度较慢
  • 适合研究认知科学与智能决策的学者

人类在认知资源有限的情况下仍能高效规划,可能得益于一套可适应的规划策略及其使用时机的掌握。然而这些策略是如何形成的?以往研究多关注个体如何选择已有策略,却很少探讨新策略的生成过程。本文提出,新策略可通过元认知强化学习发现。我们设计了一项新实验以研究这一过程,并构建了元认知强化学习模型。结果显示,该模型具备策略发现能力,且对人类策略发现行为的解释优于其他学习机制。但当拟合真实人类数据时,模型的发现速度仍慢于人类,表明仍有改进空间。

原文摘要 · Abstract (English)

One explanation for how people can plan efficiently despite limited cognitive resources is that we possess a set of adaptive planning strategies and know when and how to use them. But how are these strategies acquired? While previous research has studied how individuals learn to choose among existing strategies, little is known about the process of forming new planning strategies. In this work, we propose that new planning strategies are discovered through metacognitive reinforcement learning. To test this, we designed a novel experiment to investigate the discovery of new planning strategies. We then present metacognitive reinforcement learning models and demonstrate their capability for strategy discovery as well as show that they provide a better explanation of human strategy discovery than alternative learning mechanisms. However, when fitted to human data, these models exhibit a slower discovery rate than humans, leaving room for improvement.

认知建模强化学习策略发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。