arXiv:2608.23000cs.ROcs.HC2026-08

让机器人实时学新动作,还能不忘旧技能。

Free-Energy-Gated Plasticity for Real-Time Online Motor Learning in Physical Human-Robot Interaction

论文配图:Free-Energy-Gated Plasticity for Real-Time Online Motor Learning in Physical Human-Robot Interaction
图 1 · 摘自论文原文
  • 根据预测误差动态调节学习速度,避免遗忘旧动作。
  • 无需预训练,随机初始化网络就能学会三种循环动作。
  • 适合需要持续学习的物理人机交互场景。

全在线具身学习要求突触适应在持续互动中获取新行为,同时保留已学动力学。我们扩展了受预测编码启发的变分递归神经网络(PV-RNN),使其连续调整突触权重,并提出自由能门控可塑性(FEGP),根据变分自由能调节有效学习率。在实时物理人机交互中,一个随机初始化的网络在无离线预训练、重放或任务边界信号的情况下,自主掌握了三种循环运动模式,且所有模式均能在自主推演中重现。在十个随机教学流和每流五个网络初始化条件下控制实验显示,FEGP显著提升了技能库覆盖范围,并在离开近期观测窗口后仍保持对先前习得模式的记忆。无论是匹配门控时间平均有效率的恒定学习率,还是以破坏时间结构的方式重放相同增益值,均无法再现这些改进。结果表明,在持续在线学习中,维持已有行为的关键不在于可塑性的平均幅度或分布,而在于其相对于模型-环境不匹配的时间分配。

原文摘要 · Abstract (English)

Fully online embodied learning requires synaptic adaptation to acquire new behaviors while preserving previously learned dynamics during ongoing interaction. We extend the Predictive-Coding-inspired Variational Recurrent Neural Network (PV-RNN) to continuously adapt its synaptic weights and propose Free-Energy-Gated Plasticity (FEGP), which regulates the effective learning rate according to variational free energy. In real-time physical human-robot interaction, a randomly initialized network acquired three cyclic motor patterns without offline pretraining, replay, or task-boundary signals, with all three patterns emerging in autonomous rollouts. Controlled experiments over ten randomized teaching streams and five network initializations per stream showed that FEGP substantially improved repertoire coverage and retention of previously acquired patterns after they left the recent observation window. Neither a constant learning rate matched to the gate's time-averaged effective rate nor replay of the same gain values with disrupted temporal organization reproduced these improvements. These results indicate that the temporal allocation of plasticity relative to model-environment mismatch, rather than simply its average magnitude or distribution, is critical for maintaining previously acquired behaviors during continued online learning.

在线学习人机交互神经网络可塑性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。