研究人类如何发现新规划策略,发现三种认知机制能提升模型表现。
Individual differences in the cognitive mechanisms of planning strategy discovery
- 引入元认知伪奖励、努力估值等机制改进强化学习模型
- 三种机制显著促进策略发现,但模型仍慢于人类
- 揭示个体差异,适合研究认知建模与人机协同的学者
人们采用高效规划策略,但这些策略如何习得?已有研究指出,人们可通过强化学习从反馈中发现新策略,这一过程称为元认知强化学习(MCRL)。尽管现有模型能解释更多人的经验驱动发现,但仍存在显著个体差异,且学习速度慢于人类。本研究探讨是否加入可促进人类策略发现的认知机制,如元认知伪奖励、主观努力估值和终止反思,能使MCRL模型更贴近人类表现。对规划任务数据的分析显示,多数参与者至少使用了一种机制,其使用频率和影响因人而异。元认知伪奖励、努力估值及放弃进一步规划的价值学习均有助于策略发现。尽管这些改进揭示了个体差异及其对策略发现的影响,模型与人类表现间的差距仍未完全弥合,提示需进一步探索人类可能使用的其他发现机制。
原文摘要 · Abstract (English)
People employ efficient planning strategies. But how are these strategies acquired? Previous research suggests that people can discover new planning strategies through learning from reinforcements, a process known as metacognitive reinforcement learning (MCRL). While prior work has shown that MCRL models can learn new planning strategies and explain more participants' experience-driven discovery better than alternative mechanisms, it also revealed significant individual differences in metacognitive learning. Furthermore, when fitted to human data, these models exhibit a slower rate of strategy discovery than humans. In this study, we investigate whether incorporating cognitive mechanisms that might facilitate human strategy discovery can bring models of MCRL closer to human performance. Specifically, we consider intrinsically generated metacognitive pseudo-rewards, subjective effort valuation, and termination deliberation. Analysis of planning task data shows that a larger proportion of participants used at least one of these mechanisms, with significant individual differences in their usage and varying impacts on strategy discovery. Metacognitive pseudo-rewards, subjective effort valuation, and learning the value of acting without further planning were found to facilitate strategy discovery. While these enhancements provided valuable insights into individual differences and the effect of these mechanisms on strategy discovery, they did not fully close the gap between model and human performance, prompting further exploration of additional factors that people might use to discover new planning strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。