arXiv:2509.04633cs.NEcs.AI2025-09中稿 · NeurIPS

用AI生成的训练课程,让脑组织学会预测环境并做出反应。

The Physical Basis of Prediction: World Model Formation in Neural Organoids via an LLM-Generated Curriculum

  • 用三个渐进式虚拟环境训练脑组织,模拟学习过程。
  • 通过电生理等手段测量突触可塑性,验证模型形成。
  • 用大模型自动设计实验,提升研究效率,适合神经科学与智能研究者。

具备感知、预测和互动能力的具身智能体,其核心依赖于内部世界模型。本文提出一种新框架,用于研究生物基质中世界模型的形成与适应:人类神经类器官。我们设计了三套可扩展的闭环虚拟环境,用于训练这些生物智能体,并探究学习背后的突触机制,如长时程增强(LTP)和长时程抑制(LTD)。三种任务环境依次要求更复杂的模型:(1)条件回避任务,学习静态状态-动作关联;(2)一维捕食者-猎物场景,实现目标导向交互;(3)经典乒乓球游戏复现,建模动态连续系统。每个环境均定义状态与动作空间、感官编码与运动解码机制,以及基于可预测奖励与不可预测惩罚的反馈协议,驱动模型优化。方法上,提出元学习策略,由大语言模型自动生成并优化实验协议,显著提升环境与课程设计的规模。最后,构建多模态评估体系,超越任务表现,直接量化电生理、细胞及分子水平的突触可塑性,揭示所学世界模型的物理基础。本工作连接了基于模型的强化学习与计算神经科学,为具身智能、决策机制与智能物理基础研究提供独特平台。

原文摘要 · Abstract (English)

The capacity of an embodied agent to understand, predict, and interact with its environment is fundamentally contingent on an internal world model. This paper introduces a novel framework for investigating the formation and adaptation of such world models within a biological substrate: human neural organoids. We present a curriculum of three scalable, closed-loop virtual environments designed to train these biological agents and probe the underlying synaptic mechanisms of learning, such as long-term potentiation (LTP) and long-term depression (LTD). We detail the design of three distinct task environments that demand progressively more sophisticated world models for successful decision-making: (1) a conditional avoidance task for learning static state-action contingencies, (2) a one-dimensional predator-prey scenario for goal-directed interaction, and (3) a replication of the classic Pong game for modeling dynamic, continuous-time systems. For each environment, we formalize the state and action spaces, the sensory encoding and motor decoding mechanisms, and the feedback protocols based on predictable (reward) and unpredictable (punishment) stimulation, which serve to drive model refinement. In a significant methodological advance, we propose a meta-learning approach where a Large Language Model automates the generative design and optimization of experimental protocols, thereby scaling the process of environment and curriculum design. Finally, we outline a multi-modal evaluation strategy that moves beyond task performance to directly measure the physical correlates of the learned world model by quantifying synaptic plasticity at electrophysiological, cellular, and molecular levels. This work bridges the gap between model-based reinforcement learning and computational neuroscience, offering a unique platform for studying embodiment, decision-making, and the physical basis of intelligence.

神经类器官世界模型大模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。