提出交互式机器人新框架,让机器更好与人协同工作。
Redefining Robot Generalization Through Interactive Intelligence
- 用神经科学启发的四模块架构实现人机实时协作。
- 强调多智能体交互,支持动态适应和预测性响应。
- 适合可穿戴机器人、远程操控等需人机共演场景。
大规模机器学习的发展催生了具备强泛化能力的基础模型,可适配多种下游任务。尽管这些模型在机器人领域前景广阔,但当前范式仍视机器人为独立决策者,仅限于自主执行操作与导航,人类参与有限。然而,现实中大量机器人系统(如假肢、外骨骼、远程操控、神经接口)属于半自主型,需持续与人类伙伴进行互动协调,挑战了单一智能体假设。本文主张,机器人基础模型必须转向交互式多智能体视角,以应对人机实时共适应的复杂性。我们提出一种通用、受神经科学启发的架构,包含四个模块:(1) 基于感觉运动整合原理的多模态感知模块;(2) 类似认知科学中联合行动框架的临时协作模型;(3) 基于运动控制内模型理论的预测世界信念模型;(4) 模仿赫布可塑性和强化学习机制的记忆/反馈模块。尽管以人机共生系统为例,其中可穿戴设备与人体生理紧密融合,该框架广泛适用于半自主或交互式机器人场景。通过突破单一智能体设计,本文强调基础模型可实现更稳健、个性化且前瞻性的性能提升。
原文摘要 · Abstract (English)
Recent advances in large-scale machine learning have produced high-capacity foundation models capable of adapting to a broad array of downstream tasks. While such models hold great promise for robotics, the prevailing paradigm still portrays robots as single, autonomous decision-makers, performing tasks like manipulation and navigation, with limited human involvement. However, a large class of real-world robotic systems, including wearable robotics (e.g., prostheses, orthoses, exoskeletons), teleoperation, and neural interfaces, are semiautonomous, and require ongoing interactive coordination with human partners, challenging single-agent assumptions. In this position paper, we argue that robot foundation models must evolve to an interactive multi-agent perspective in order to handle the complexities of real-time human-robot co-adaptation. We propose a generalizable, neuroscience-inspired architecture encompassing four modules: (1) a multimodal sensing module informed by sensorimotor integration principles, (2) an ad-hoc teamwork model reminiscent of joint-action frameworks in cognitive science, (3) a predictive world belief model grounded in internal model theories of motor control, and (4) a memory/feedback mechanism that echoes concepts of Hebbian and reinforcement-based plasticity. Although illustrated through the lens of cyborg systems, where wearable devices and human physiology are inseparably intertwined, the proposed framework is broadly applicable to robots operating in semi-autonomous or interactive contexts. By moving beyond single-agent designs, our position emphasizes how foundation models in robotics can achieve a more robust, personalized, and anticipatory level of performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。