用单臂指令生成双臂动作,无需双臂示范即可高效完成复杂操作。
One-to-Two Acting: A Novel Framework for Single-arm Agent Action Expansion to Dual Arms

- 通过文本分解任务并建模时序依赖,自动生成可执行的双臂动作序列。
- 仿真中减少54.4%执行步数,成功率与单臂基线相当。
- 仅需少量单臂样本,适合缺乏双臂数据的机器人应用。
双臂操作可通过并行执行提升效率,但收集双臂示范数据成本高且困难。本文提出ExS2D,一种分层动作扩展框架,可基于单臂监督实现双臂操作。ExS2D首先从文本指令中生成带时序优先关系的结构化子任务;然后通过观察引导的子任务映射,将每个子任务转化为可执行动作;最后由多模态大语言模型驱动的协调器,进行优先级感知的动作分配与同步规划,选择无碰撞的双臂执行方案。仿真实验表明,ExS2D平均执行步数减少54.4%,同时保持与单臂基线相当的成功率。在四个真实任务上的机器人实验进一步验证了该方法在少样本单臂示范下实现可靠双臂执行的能力,且全程无需双臂示范。
原文摘要 · Abstract (English)
Dual-arm manipulation can improve throughput via parallel execution, but collecting bimanual demonstrations for training is costly and difficult. We present ExS2D, a hierarchical action expansion framework that enables dual-arm manipulation from single-arm supervision. ExS2D first generates structured subtasks from textual instructions while explicitly capturing temporal precedence. It then grounds each subtask into executable actions through subtask-guided action mapping in observation. Finally, precedence-aware action allocation and synchronized planning are performed by a multimodal large language model driven coordinator to select collision-free dual-arm executions. Simulation experiments demonstrate that ExS2D reduces the average execution steps by 54.4% while maintaining a comparable success rate to a single-arm baseline. Real-robot experiments on four tasks further demonstrate the reliability of ExS2D for dual-arm execution under few-shot single-arm samples, while using zero bimanual demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。