用多智能体框架让大模型零样本控制双臂机器人完成复杂操作。
Bimanual Robot Manipulation via Multi-Agent In-Context Learning

- 将双臂协作拆解为分步的主从预测,降低动作空间复杂度。
- 在TWIN基准上平均成功率70.5%,比现有无训练方法高6.1个百分点。
- 无需硬件微调,真实场景中3个任务表现优异,适合快速部署。
语言模型(LLMs)已成为具身控制的强大推理引擎。特别是上下文学习(ICL)使现成的纯文本大模型能在不进行任务特异性训练的情况下预测机器人动作,同时保持其泛化能力。然而,将ICL应用于双臂操作仍具挑战性,因其高维关节动作空间和紧密的双臂协调约束会迅速超出标准上下文窗口。为此,我们提出BiCICLe(双臂协同上下文学习),首个无需微调即可实现少样本双臂操作的标准大模型框架。BiCICLe将双臂控制建模为多智能体主从问题,将动作空间解耦为顺序的、条件化的单臂预测。在TWIN基准的13项任务上评估,BiCICLe达到70.5%的平均成功率,优于最佳无训练基线6.1个百分点,并超越多数监督方法。我们还在3个真实任务中展示了优越性能,且无需硬件特定再训练。
原文摘要 · Abstract (English)
Language Models (LLMs) have emerged as powerful reasoning engines for embodied control. In particular, In-Context Learning (ICL) enables off-the-shelf, text-only LLMs to predict robot actions without any task-specific training while preserving their generalization capabilities. Applying ICL to bimanual manipulation remains challenging as the high-dimensional joint action space and tight inter-arm coordination constraints rapidly overwhelm standard context windows. To address this, we introduce BiCICLe (Bimanual Coordinated In-Context Learning), the first framework that enables standard LLMs to perform few-shot bimanual manipulation without fine-tuning. BiCICLe frames bimanual control as a multi-agent leader-follower problem, decoupling the action space into sequential, conditioned single-arm predictions. Evaluated on 13 tasks from the TWIN benchmark, BiCICLe achieves 70.5% average success rate, outperforming the best training-free baseline by 6.1 percentage points and surpassing most supervised methods. We also demonstrate superior real-world performance on 3 tasks without hardware-specific retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。