arXiv:2509.06810q-bio.NCcs.LG2025-09被引 2

压缩奖励函数提升目标导向学习效率

Reward function compression facilitates goal-dependent reinforcement learning

  • 用压缩的奖励规则替代复杂目标,减轻工作记忆负担
  • 目标空间越适合压缩,学习速度越快,反应越敏捷
  • 适合研究认知机制与激励设计的学者参考

人类能为新颖抽象的目标赋予价值以支持强化学习,但这种灵活性代价高昂且降低学习效率。我们提出,目标导向学习初期依赖容量有限的工作记忆;通过持续经验,学习者构建出‘压缩的奖励函数’——一种简化的目标规则——转移至长期记忆,实现反馈后的自动评估。这种自动化释放了工作记忆资源,从而提升学习效率。六个实验表明,目标空间越大,学习受损越严重,而具备压缩潜力的目标空间则促进学习。更快的奖励处理速度与更好学习表现相关。尽管算法细节尚待明确,行为数据与计算模型均表明,高效的目标导向学习依赖于将复杂目标信息压缩为稳定奖励函数。这些发现揭示了内在动机的认知机制,可为人类目标达成的干预策略提供依据。

原文摘要 · Abstract (English)

Humans can uniquely assign value to novel, abstract outcomes to support reinforcement learning. However, this flexibility is cognitively costly and reduces learning efficiency. We propose that goal-dependent learning initially relies on capacity-limited working memory. With consistent experience, learners create a "compressed" reward function - a simplified goal rule -- that transfers to long-term memory for a more automatic evaluation upon receiving feedback. This automaticity frees working memory resources, thereby boosting learning efficiency. Across six experiments, we demonstrate that learning is impaired by the size of the goal space but improves when this space allows for compression. Additionally, faster reward processing correlates with better learning. Although the algorithmic details remain to be established, our behavioral results and computational models suggest that efficient goal-directed learning relies on compressing complex goal information into a stable reward function. These findings illuminate the cognitive mechanisms of intrinsic motivation and can inform behavioral interventions supporting human goal achievement.

强化学习认知机制奖励压缩目标导向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。