探索行为让预测模型更有序,揭示了行为如何塑造大脑的内在表征。
Exploratory Experience Shapes the Geometry of Predictive Representations

- 用可调节探索/利用比例的智能体在树状迷宫中学习,通过预测编码更新内部模型。
- 探索型智能体的表征更空间有序,能更好保留迷宫转移结构;利用型则较混乱。
- 小鼠实验显示,探索多的个体其神经表征与探索型智能体相似,支持该机制跨物种通用。
主动感知通过动作-感知循环连接行为与学习:动作决定感知输入,用于更新内部预测模型,进而指导下一步动作。预测编码框架自然适用于建模此过程,因内部表征持续更新以预测未来观测。本文研究探索性与利用性行为策略如何塑造这些内部预测表征。我们在树状迷宫中构建一个在线学习智能体,其行为受可调参数控制,平衡探索与利用。智能体基于自身行为生成的经验,用预测编码方法更新感知模型,该模型可预测未来迷宫状态及奖励概率,从而在探索时依据预期信息增益、在利用时依据预测奖励选择动作。结果显示,探索型智能体发展出更空间有序的表征,且在潜在空间中更好地保留迷宫转移结构;而利用型智能体则形成较无序的表征。随后,我们将该模型应用于脱水小鼠在相同迷宫中的自然轨迹数据,比较其表征与智能体轨迹的差异:探索性强的小鼠表现出与探索型智能体相似的表征几何,而访问模式受限的小鼠则类似奖励驱动的利用型智能体。结果表明,探索有助于预测模型形成泛化表征,使潜在空间围绕空间位置和转移上下文组织,这一机制在人工智能与动物中均成立。
原文摘要 · Abstract (English)
Active sensing links behavior and learning through an action-perception loop: actions determine the observations used to update internal predictive models of perception, which subsequently guide the next actions. Predictive-coding frameworks provide a natural way to model this process, since internal representations are continuously updated to predict future observations. Here, we ask how exploratory and exploitative behavioral strategies shape these internal predictive representations. We build an online learning agent in a tree-like maze with a controllable parameter regulating the balance between exploratory and exploitative regimes. The agent updates a predictive-coding-based perception model from experience generated by its own behavior. The model predicts both future maze states and reward probability, allowing the agent to select actions either by expected information gain during exploration or by predicted reward during exploitation. We show that the resulting internal predictive representations depend strongly on the agent's behavioral regime. Exploratory agents develop representations that are more spatially organized and better preserve the structure of maze transitions in latent space. In contrast, exploitative agents learn less organized representations. We then train this predictive model on natural trajectories of water-deprived mice navigating the same maze and compare the resulting representations with those learned from agent trajectories. More exploratory mice show representational geometries that closely match those of exploratory agents, whereas mice with more restricted visitation patterns resemble reward-driven, exploitative agents. Together, these findings suggest that exploration enables predictive models to form generalized internal representations by organizing latent space around both spatial location and transition context in artificial agents and animals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。