提出实时协作框架MonTA,让机器人更快响应人类伙伴。
Towards Real-time Adaptation of Embodied Agent in Human-Robot Collaboration
- 分层设计:高频监控+低频适配,兼顾速度与推理
- 在多种协作场景中表现更优,任务完成率显著提升
- 适合需快速响应的机器人协作研究与应用
大语言模型为人机协作带来新可能,但多数模型延迟高,难以实现实时协作。为此,我们首先提出一个细粒度基准,专门评估代理在Overcooked-AI环境中的主动适应能力和时间响应性。基于评估结果,我们提出受认知科学启发的分层框架MonTA(Monitor-then-Adapt),包含三个模块:以7 Hz高频运行的轻量级监控器,用于检测适应需求;以及两个用于子任务和路径适应推理的高效适配器,以较低频率向人类提供指令。实验表明,MonTA在所提基准上显著优于基线模型,在不同团队协作流畅度布局下均表现更优。用户研究证实,该框架生成的适应方案合理、语言指令一致可靠。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have opened transformative possibilities for human-robot collaboration. However, enabling real-time collaboration requires both low latency and robust reasoning, and most LLMs suffer from high latency. To address this gap, we first propose a fine-grained benchmark that explicitly assesses agents' proactive adaptability and temporal responsiveness in the Overcooked-AI environment. Based on evaluation results, we propose MonTA (Monitor-then-Adapt), a hierarchical framework inspired by cognitive science research. MonTA contains three key modules: a lightweight Monitor that operates at high frequency (7 Hz) to detect adaptation needs, and two proficient Adapters for subtask and path adaptation reasoning that provide instructions to humans at a lower frequency. Our results demonstrate that MonTA significantly outperforms baseline agents on our proposed benchmark, achieving superior performance across layouts with varying teaming fluency. User studies confirm the high reasonableness of adaptation plans and consistent language instructions provided by our framework to humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。