用脑电波信号加速机器人高维操作任务的强化学习。
Accelerating Reinforcement Learning via Error-Related Human Brain Signals
- 从脑电图中解码错误相关电位,用于奖励塑形。
- 在障碍物环境中,成功率达78%,超越稀疏奖励基线。
- 对不同用户鲁棒,适合人机协同技能学习场景。
本文研究如何利用隐式神经反馈加速复杂机器人操作任务中的强化学习。以往基于脑电图(EEG)的强化学习多聚焦于导航或低维运动任务,而本文探讨其在高维操作任务(含障碍物与末端执行器精控)中的适用性。通过离线训练的EEG分类器解码错误相关电位,并融入奖励塑形,系统评估人类反馈权重的影响。在7自由度机械臂的障碍密集抓取环境中实验显示,神经反馈可显著加速学习,最优反馈权重下任务成功率最高达78%,部分情形超过稀疏奖励基线。跨被试分析表明,最佳权重设定下仍具一致性加速效果;剔除单个被试的留一法验证进一步证实框架对个体间EEG解码差异的鲁棒性。结果表明,基于EEG的强化学习可拓展至非移动类任务,为实现人机对齐的操作技能习得提供可行路径。
原文摘要 · Abstract (English)
In this work, we investigate how implicit neural feed back can accelerate reinforcement learning in complex robotic manipulation settings. While prior electroencephalogram (EEG) guided reinforcement learning studies have primarily focused on navigation or low-dimensional locomotion tasks, we aim to understand whether such neural evaluative signals can improve policy learning in high-dimensional manipulation tasks involving obstacles and precise end-effector control. We integrate error related potentials decoded from offline-trained EEG classifiers into reward shaping and systematically evaluate the impact of human-feedback weighting. Experiments on a 7-DoF manipulator in an obstacle-rich reaching environment show that neural feedback accelerates reinforcement learning and, depending on the human-feedback weighting, can yield task success rates that at times exceed those of sparse-reward baselines. Moreover, when applying the best-performing feedback weighting across all sub jects, we observe consistent acceleration of reinforcement learning relative to the sparse-reward setting. Furthermore, leave-one subject-out evaluations confirm that the proposed framework remains robust despite the intrinsic inter-individual variability in EEG decodability. Our findings demonstrate that EEG-based reinforcement learning can scale beyond locomotion tasks and provide a viable pathway for human-aligned manipulation skill acquisition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。