用自驱强化学习让AI自动发现并操控复杂系统中的自组织现象。
The Artificial Experimentalist: Discovery and Control of Self-Organizing Phenomena with Autotelic Reinforcement Learning

- AI在模拟中自主设定目标,通过局部扰动干预系统演化。
- 发现稳定准粒子的速度比传统方法快,还能控制其运动方向。
- 可零样本泛化到新规则和新场景,适合人机协同实验设计。
现有探索元胞自动机等复杂系统的方法多为开环:设定初始条件后执行完整模拟并观察结果,过程中不进行干预。本文提出基于自驱强化学习的闭环框架,使智能体自主采样多样目标,并学习目标条件策略,以最小、局部扰动干预复杂系统。我们在以生命般自组织模式著称的连续元胞自动机Lenia上实现该框架,构建名为CARL的代理系统,展示了三项能力:第一,CARL在多种Lenia更新规则下发现稳定孤子的速度高于启发式基线;第二,它能仅用少量干预即控制已有孤子的运动方向,证明其具备对自组织现象的调控能力;第三,人类可通过指定高层方向指令,实时引导训练好的代理将孤子穿越迷宫环境。代理在多样化目标、更新规则和随机初始状态下训练,获得可零样本泛化至各类分布外条件的策略。这些结果表明,通往自主或人机协作的‘人工实验学家’代理之路已现曙光。
原文摘要 · Abstract (English)
Existing methods for exploring cellular automata and other complex systems mostly operate in open loop: they set initial conditions, execute a full simulation, and observe the outcome, without intervening during execution. We introduce a closed-loop framework based on autotelic reinforcement learning, in which an agent autonomously samples diverse goals and learns a goal-conditioned policy to intervene in a complex system through minimal, local perturbations. We instantiate this framework on Lenia, a continuous cellular automaton known for life-like self-organizing patterns, in an agentic system we call CARL, and demonstrate three capabilities. First, CARL discovers stable solitons across a wide range of Lenia update rules at a higher rate than heuristic baselines. Second, it learns to steer the movement direction of existing solitons with few interventions, showing that CARL can control self-organizing patterns, not only create them. Third, humans can use trained agents to guide solitons through maze environments in real time by specifying high-level directional commands that the agent translates into low-level interventions. Trained across diverse goals, update rules, and random initial states, the agents acquire policies that generalize zero-shot to various out-of-distribution conditions. These results suggest a path toward artificial experimentalist agents that, autonomously or with human guidance, discover and control emergent phenomena in complex systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。