让人形机器人与人类协作搬运,实现长期规划与实时控制的无缝衔接。
Cognition to Control - Multi-Agent Learning for Human-Humanoid Collaborative Transport
- 分三层架构:感知、决策、执行,明确从思考到动作的路径。
- 在协作任务中成功率更高,能自动生成领导者和跟随者角色。
- 适合需要长时间协调的多智能体人机协作场景。
高效的人-机器人协作(HRC)需将高层意图转化为稳定接触的全身运动,并持续适应人类伙伴。现有视觉-语言-动作(VLA)系统多采用端到端映射,但侧重反应式行为,缺乏如何融入持续性系统2式推理与可靠低延迟连续控制的机制。这一差距在多智能体HRC中尤为突出,因长期协调决策与物理执行必须在接触、可行性与安全约束下协同演化。本文提出认知到控制(C2C)三层次框架:(i) 基于VLM的感知层,保持场景指代一致性并推断具身化可操作性/约束;(ii) 决策层——系统2核心,通过去中心化多智能体强化学习(MARL)建模为共享势能函数的马尔可夫势博弈,优化长期技能选择与序列;(iii) 全身控制层以高频执行选定技能,保障运动学/动力学可行性与接触稳定性。决策层以残差策略形式实现,相对于基准控制器,内化伙伴动态且无需显式角色分配。在协作操作任务上的实验表明,其成功率与鲁棒性优于单智能体及端到端基线,具备稳定协调能力并涌现领导者-跟随者行为。
原文摘要 · Abstract (English)
Effective human-robot collaboration (HRC) requires translating high-level intent into contact-stable whole-body motion while continuously adapting to a human partner. Many vision-language-action (VLA) systems learn end-to-end mappings from observations and instructions to actions, but they often emphasize reactive (System 1-like) behavior and leave under-specified how sustained System 2-style deliberation can be integrated with reliable, low-latency continuous control. This gap is acute in multi-agent HRC, where long-horizon coordination decisions and physical execution must co-evolve under contact, feasibility, and safety constraints. We address this limitation with cognition-to-control (C2C), a three-layer hierarchy that makes the deliberation-to-control pathway explicit: (i) a VLM-based grounding layer that maintains persistent scene referents and infers embodiment-aware affordances/constraints; (ii) a deliberative skill/coordination layer-the System 2 core-that optimizes long-horizon skill choices and sequences under human-robot coupling via decentralized MARL cast as a Markov potential game with a shared potential encoding task progress; and (iii) a whole-body control layer that executes the selected skills at high frequency while enforcing kinematic/dynamic feasibility and contact stability. The deliberative layer is realized as a residual policy relative to a nominal controller, internalizing partner dynamics without explicit role assignment. Experiments on collaborative manipulation tasks show higher success and robustness than single-agent and end-to-end baselines, with stable coordination and emergent leader-follower behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。