用脑电波信号实现机器人智能控制,无需用户手动评分。
Reinforcement Learning from Implicit Neural Feedback for Human-Aligned Robot Control
- 通过脑电图中的错误信号提取隐式反馈,替代传统按键评分
- 在稀疏奖励下仍达到与密集奖励相当的控制性能
- 适合人机交互、无障碍控制等需要自然协作的场景
传统强化学习在稀疏奖励下难以有效学习,常需人工设计复杂奖励函数。为此,基于人类反馈的强化学习(RLHF)应运而生,但多数方法依赖按钮或偏好标签等显式反馈,干扰自然交互并增加认知负担。本文提出一种基于隐式人类反馈的强化学习(RLIHF)框架,利用非侵入式脑电图(EEG)信号中的错误相关电位(ErrPs)提供连续、无干预的反馈。该方法采用预训练解码器将原始脑电信号转换为概率奖励分量,即使在外部奖励稀疏的情况下也能实现有效策略学习。我们在基于MuJoCo物理引擎的仿真环境中,使用Kinova Gen2机械臂完成避障抓取任务进行评估。结果表明,使用解码脑电反馈训练的智能体性能可媲美使用密集人工奖励训练的模型。研究验证了利用神经隐式反馈实现可扩展、以人为本的强化学习在交互式机器人中的可行性。
原文摘要 · Abstract (English)
Conventional reinforcement learning (RL) approaches often struggle to learn effective policies under sparse reward conditions, necessitating the manual design of complex, task-specific reward functions. To address this limitation, reinforcement learning from human feedback (RLHF) has emerged as a promising strategy that complements hand-crafted rewards with human-derived evaluation signals. However, most existing RLHF methods depend on explicit feedback mechanisms such as button presses or preference labels, which disrupt the natural interaction process and impose a substantial cognitive load on the user. We propose a novel reinforcement learning from implicit human feedback (RLIHF) framework that utilizes non-invasive electroencephalography (EEG) signals, specifically error-related potentials (ErrPs), to provide continuous, implicit feedback without requiring explicit user intervention. The proposed method adopts a pre-trained decoder to transform raw EEG signals into probabilistic reward components, enabling effective policy learning even in the presence of sparse external rewards. We evaluate our approach in a simulation environment built on the MuJoCo physics engine, using a Kinova Gen2 robotic arm to perform a complex pick-and-place task that requires avoiding obstacles while manipulating target objects. The results show that agents trained with decoded EEG feedback achieve performance comparable to those trained with dense, manually designed rewards. These findings validate the potential of using implicit neural feedback for scalable and human-aligned reinforcement learning in interactive robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。