用脑电接口让机器人感知人何时忙,自动延迟打扰。
Multi-Modal Multi-Agent Robotic Cognitive Alignment enabled by Non-Invasive Consumer Brain Computer Interfaces: A Proof of Concept Exploration
- 用消费级脑机接口实时监测脑电波,判断人类专注程度。
- 高专注时暂停主代理消息,由副代理后台处理任务。
- 适合需要无感交互的智能机器人系统设计者参考。
尽管非语言行为和表达性动作对自然的人机交互至关重要,现有方法常忽视一个关键因素:人类的内部认知状态。主动式多智能体系统常在不恰当时机打断用户,导致认知过载并降低任务表现。本文提出一种生成“认知对齐”多智能体交互的框架,提升机器人系统在人类高心理负荷与高参与度时刻,上下文感知地推迟通信的能力。我们设计并实现了一个闭环架构,探索自主任务执行与实时神经生理注意力之间的相互作用。通过消费级脑机接口(BCI),在人类执行高参与度任务时持续监测脑电图(EEG)频段功率。提出一种基于参与度的处理流程:当检测到高参与度时,通过基于HTTP的信号机制将主代理的感官输入与音频输出置于等待状态,使次级代理能无缝在后台处理复杂任务。一旦人类认知状态回归低负荷基线,主代理释放已排队的消息。初步结果表明,利用实时信号处理、大语言模型(LLMs)和物理机器人实体,可实现认知感知、非侵入式的多智能体系统。
原文摘要 · Abstract (English)
While non-verbal behaviors and expressive movements are essential for natural human-robot interaction, existing methods often overlook a crucial element: the human's internal cognitive state. Frequently, proactive multi-agent systems can interrupt humans at inopportune moments, leading to cognitive overload and decreased task performance. This paper introduces a framework for generating "cognitively aligned" multi-agent interactions, enhancing the ability of robotic systems to contextually defer communications to the user of an agent system during moments of high human mental workload and engagement. We present the design and implementation of a closed-loop architecture that explores the interplay between autonomous task execution and real-time neurophysiological focus. Using a consumer-grade Brain-Computer Interface (BCI), our approach continuously monitors Electroencephalography (EEG) spectral band powers while a human performs an engagement-inducing task. We propose an engagement-driven pipeline where an HTTP-based signaling mechanism places a primary agent's sensory inputs and audio outputs into a holding state upon detecting high engagement. This allows secondary agents to seamlessly process complex, delegated tasks in the background. Once the human's cognitive state returns to a lower cognitive load baseline, the primary agent releases the queued agent message. Our preliminary results demonstrate the feasibility of leveraging real-time signal processing, Large Language Models (LLMs), and physical robotic embodiments to create cognitively-aware, non-intrusive multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。