arXiv:2511.05275cs.ROcs.LG2025-11中稿 · ICLR被引 6

用两个单臂模型组合实现高效双臂操作,无需额外双臂数据。

TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models

  • 将预训练单臂视觉语言动作模型复制两份,协同完成双臂任务。
  • 在真实与仿真场景中超越同规模单体模型,接近顶尖水平。
  • 仅依赖公开单臂数据,适合资源有限的研究者快速复现。

视觉语言动作模型(VLAs)在大规模机器人数据集上训练后,在操作任务中表现出色,包括双臂任务。然而,由于多数公开数据集聚焦于单臂示范,将VLAs用于双臂任务通常需要大量额外的双臂数据和微调。为此,我们提出TwinVLA,一种模块化框架,将两个预训练的单臂VLA组合成协调的双臂VLA。不同于在单臂与双臂数据混合上训练的单体跨体态模型,TwinVLA通过组合预训练单臂策略,提升了数据效率与性能。在真实世界与仿真环境中的多种双臂任务中,TwinVLA的表现优于同等规模的单体RDT-1B模型,且无需任何双臂预训练。此外,其性能逼近依赖大量专有双臂数据与高算力的顶级模型π₀。这些结果证明,模块化组合方法是利用公开单臂数据实现高性能双臂操作的高效且可扩展路径。

原文摘要 · Abstract (English)

Vision-language-action models (VLAs) trained on large-scale robotic datasets have demonstrated strong performance on manipulation tasks, including bimanual tasks. However, because most public datasets focus on single-arm demonstrations, adapting VLAs for bimanual tasks typically requires substantial additional bimanual data and fine-tuning. To address this challenge, we introduce TwinVLA, a modular framework that composes two copies of a pretrained single-arm VLA into a coordinated bimanual VLA. Unlike monolithic cross-embodiment models trained on mixtures of single-arm and bimanual data, TwinVLA improves both data efficiency and performance by composing pretrained single-arm policies. Across diverse bimanual tasks in real-world and simulation settings, TwinVLA outperforms a comparably-sized monolithic RDT-1B model without requiring any bimanual pretraining. Furthermore, it narrows the gap to state-of-the-art model $π_0$, which relies on extensive proprietary bimanual data and compute cost. These results establish our modular composition approach as a data-efficient and scalable path toward high-performance bimanual manipulation, leveraging public single-arm data.

双臂操作视觉语言动作数据效率模块化设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。