arXiv:2603.15248cs.LG2026-03

揭示婴儿运动学习中目标控制的机制原理,解释策略如何演化并竞争。

Mechanistic Foundations of Goal-Directed Control

  • 基于感应-运动-认知系统,构建可解释的控制电路模型。
  • 发现当上下文窗口k≥8时,门控信心随log k渐近增长。
  • 适合研究认知发展与可解释智能体设计的学者参考。

机制可解释性已推动Transformer电路分析,分解行为为竞争算法、识别训练中的相变,并推导策略切换的闭式预测。然而该框架仍局限于序列预测架构,缺乏对具身控制系统机制性解释。本文将此框架扩展至感知-运动-认知发展,以婴儿运动学习为模型系统。结果表明,基础归纳偏置催生因果控制电路,学习到的门控机制收敛于理论上的不确定性阈值。动态分析显示仲裁门存在清晰相变,其承诺行为可由闭式指数移动平均代理模型精确描述。关键参数上下文窗口k决定电路形成:当k≤4时无法形成仲裁机制;当k≥8时,门控信心渐近增长,符合log k规律。二维相图进一步揭示任务需求依赖的路径仲裁,支持前瞻性执行仅在预测误差低于任务容差窗口时才具优势的理论。上述成果提供了学习过程中反应式与前瞻性控制策略生成与竞争的机制解释。更广泛地,本工作提升了认知发展的机制理解,并为可解释具身智能体的设计提供原则性指导。

原文摘要 · Abstract (English)

Mechanistic interpretability has transformed the analysis of transformer circuits by decomposing model behavior into competing algorithms, identifying phase transitions during training, and deriving closed-form predictions for when and why strategies shift. However, this program has remained largely confined to sequence-prediction architectures, leaving embodied control systems without comparable mechanistic accounts. Here we extend this framework to sensorimotor-cognitive development, using infant motor learning as a model system. We show that foundational inductive biases give rise to causal control circuits, with learned gating mechanisms converging toward theoretically motivated uncertainty thresholds. The resulting dynamics reveal a clean phase transition in the arbitration gate whose commitment behavior is well described by a closed-form exponential moving-average surrogate. We identify context window k as the critical parameter governing circuit formation: below a minimum threshold (k$\leq$4) the arbitration mechanism cannot form; above it (k$\geq$8), gate confidence scales asymptotically as log k. A two-dimensional phase diagram further reveals task-demand-dependent route arbitration consistent with the prediction that prospective execution becomes advantageous only when prediction error remains within the task tolerance window. Together, these results provide a mechanistic account of how reactive and prospective control strategies emerge and compete during learning. More broadly, this work sharpens mechanistic accounts of cognitive development and provides principled guidance for the design of interpretable embodied agents.

机制可解释认知发展具身智能控制策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。