用脑电波信号实现人机对齐,让机器人自动学习复杂任务。
Aligning Humans and Robots via Reinforcement Learning from Implicit Human Feedback
- 通过脑电波中的错误信号获取隐式反馈,无需用户主动操作。
- 在模拟环境中,性能接近人工密集奖励训练的水平。
- 适合人机交互、神经接口与智能机器人研究者。
传统强化学习在稀疏奖励下表现不佳,需人工设计复杂奖励函数。为克服此问题,我们提出一种基于隐式人类反馈的强化学习框架(RLIHF),利用非侵入式脑电图(EEG)信号中的错误相关电位(ErrPs)提供连续、无干扰的反馈。该方法采用预训练解码器将原始脑电信号转换为概率奖励分量,使智能体在外部奖励稀疏时仍能有效学习。我们在基于MuJoCo物理引擎的仿真环境中,使用Kinova Gen2机械臂完成避障抓取任务进行评估。结果表明,使用解码脑电反馈训练的智能体性能可媲美采用密集人工奖励训练的模型,验证了隐式神经反馈在可扩展、以人为本的机器人强化学习中的潜力。
原文摘要 · Abstract (English)
Conventional reinforcement learning (RL) ap proaches often struggle to learn effective policies under sparse reward conditions, necessitating the manual design of complex, task-specific reward functions. To address this limitation, rein forcement learning from human feedback (RLHF) has emerged as a promising strategy that complements hand-crafted rewards with human-derived evaluation signals. However, most existing RLHF methods depend on explicit feedback mechanisms such as button presses or preference labels, which disrupt the natural interaction process and impose a substantial cognitive load on the user. We propose a novel reinforcement learning from implicit human feedback (RLIHF) framework that utilizes non-invasive electroencephalography (EEG) signals, specifically error-related potentials (ErrPs), to provide continuous, implicit feedback without requiring explicit user intervention. The proposed method adopts a pre-trained decoder to transform raw EEG signals into probabilistic reward components, en abling effective policy learning even in the presence of sparse external rewards. We evaluate our approach in a simulation environment built on the MuJoCo physics engine, using a Kinova Gen2 robotic arm to perform a complex pick-and-place task that requires avoiding obstacles while manipulating target objects. The results show that agents trained with decoded EEG feedback achieve performance comparable to those trained with dense, manually designed rewards. These findings validate the potential of using implicit neural feedback for scalable and human-aligned reinforcement learning in interactive robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。