用神经网络从行为数据中灵活推断动物学习规则,比传统方法更准确。
Flexible inference for animal learning rules using neural networks
- 用深度神经网络建模学习规则,直接从行为数据中推断。
- 在小鼠任务学习数据上,预测表现优于传统强化学习模型。
- 能捕捉奖励历史依赖的复杂学习动态,适合行为建模研究者。
理解动物如何学习是神经科学的核心挑战,对发展与动物或人类对齐的人工智能具有重要意义。然而,现有方法通常假设学习规则具有固定参数形式(如Q-learning、策略梯度),可能无法准确描述动物在真实场景中的复杂学习行为。为此,我们提出一种框架,直接从新任务学习过程中的行为数据中推断学习规则。假设动物的决策策略由广义线性模型(GLM)参数化,并使用深度神经网络(DNN)建模其学习规则——即从任务协变量到每轮权重更新的映射。该方法实现灵活的数据驱动推断,同时保持决策策略的可解释性。为捕捉更复杂的动态,引入循环神经网络(RNN)变体,放宽学习仅依赖当前试验协变量的马尔可夫假设,允许整合多轮信息。模拟结果表明框架可恢复真实学习规则。我们将DNN和RNN方法应用于大样本小鼠感官决策任务行为数据,发现其在预测保留小鼠的学习轨迹方面优于传统强化学习规则。推断出的学习规则表现出依赖奖励历史的学习动态,连续受奖后更新幅度更大。这些方法为新任务学习中从行为数据推断学习规则提供了灵活框架,有助于改进动物训练方案并推动行为数字孪生的发展。
原文摘要 · Abstract (English)
Understanding how animals learn is a central challenge in neuroscience, with growing relevance to the development of animal- or human-aligned artificial intelligence. However, existing approaches tend to assume fixed parametric forms for the learning rule (e.g., Q-learning, policy gradient), which may not accurately describe the complex forms of learning employed by animals in realistic settings. Here we address this gap by developing a framework to infer learning rules directly from behavioral data collected during de novo task learning. We assume that animals follow a decision policy parameterized by a generalized linear model (GLM), and we model their learning rule -- the mapping from task covariates to per-trial weight updates -- using a deep neural network (DNN). This formulation allows flexible, data-driven inference of learning rules while maintaining an interpretable form of the decision policy itself. To capture more complex learning dynamics, we introduce a recurrent neural network (RNN) variant that relaxes the Markovian assumption that learning depends solely on covariates of the current trial, allowing for learning rules that integrate information over multiple trials. Simulations demonstrate that the framework can recover ground-truth learning rules. We applied our DNN and RNN-based methods to a large behavioral dataset from mice learning to perform a sensory decision-making task and found that they outperformed traditional RL learning rules at predicting the learning trajectories of held-out mice. The inferred learning rules exhibited reward-history-dependent learning dynamics, with larger updates following sequences of rewarded trials. Overall, these methods provide a flexible framework for inferring learning rules from behavioral data in de novo learning tasks, setting the stage for improved animal training protocols and the development of behavioral digital twins.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。