arXiv:2501.12633cs.LGcs.AI2025-01ICML被引 9

用历史依赖奖励建模动物长期决策,更真实捕捉自然行为。

Inverse Reinforcement Learning with Switching Rewards and History Dependency for Characterizing Animal Behaviors

  • 引入时变历史依赖奖励函数,模拟动物基于过往的决策机制
  • 在模拟与真实数据上均优于无历史依赖的模型,提升预测准确率
  • 适合研究自然状态下动物内在动机与复杂行为的科研人员

传统神经科学中的决策研究多聚焦于简单重复的任务,动物通过明确奖励获得短期行为反馈。然而,这限制了对由内在动机驱动、长期且复杂的自然行为的理解。近期时变逆强化学习(IRL)尝试捕捉长期自由行为中的动机变化,但忽略了动物决策依赖历史事实这一关键特征。为此,本文提出SWIRL(SWitching IRL)框架,将长期行为序列建模为多个短时决策过程的切换,每个过程由独立的时变奖励函数控制,并引入生物学合理的记忆依赖机制,以反映过去决策与环境背景对当前行为的影响。在模拟和真实动物行为数据集上的实验表明,该方法在定量与定性层面均显著优于缺乏历史依赖性的模型。这是首个同时融合历史依赖策略与奖励的逆强化学习模型,推动了对动物复杂自然决策机制的理解。

原文摘要 · Abstract (English)

Traditional approaches to studying decision-making in neuroscience focus on simplified behavioral tasks where animals perform repetitive, stereotyped actions to receive explicit rewards. While informative, these methods constrain our understanding of decision-making to short timescale behaviors driven by explicit goals. In natural environments, animals exhibit more complex, long-term behaviors driven by intrinsic motivations that are often unobservable. Recent works in time-varying inverse reinforcement learning (IRL) aim to capture shifting motivations in long-term, freely moving behaviors. However, a crucial challenge remains: animals make decisions based on their history, not just their current state. To address this, we introduce SWIRL (SWitching IRL), a novel framework that extends traditional IRL by incorporating time-varying, history-dependent reward functions. SWIRL models long behavioral sequences as transitions between short-term decision-making processes, each governed by a unique reward function. SWIRL incorporates biologically plausible history dependency to capture how past decisions and environmental contexts shape behavior, offering a more accurate description of animal decision-making. We apply SWIRL to simulated and real-world animal behavior datasets and show that it outperforms models lacking history dependency, both quantitatively and qualitatively. This work presents the first IRL model to incorporate history-dependent policies and rewards to advance our understanding of complex, naturalistic decision-making in animals.

逆强化学习动物行为历史依赖决策建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。