arXiv:2603.03836cs.RO2026-03被引 3

通过技能复用提升双臂操作的组合多样性,解决传统模型无法灵活搭配动作的问题。

SkillVLA: Tackling Combinatorial Diversity in Dual-Arm Manipulation via Skill Reuse

  • 分离左右臂技能表示,支持已学单臂动作自由组合
  • 在新配对下成功率从0%提升至51%
  • 适合需要协同与长序列动作的复杂操作任务

视觉-语言-动作(VLA)模型在双臂操作中展现出强大潜力,能实现复杂行为并泛化到未见环境。然而,主流双臂VLA方法普遍忽视了组合多样性这一关键挑战:不同单臂行为的配对会产生质变的任务行为,但现有模型未显式建模此结构。我们提出,高效的双臂VLA应支持技能复用——即重新组合先前学习的单臂技能,以应对新的左右臂配对,避免为每种组合单独训练。当前VLA设计将技能耦合于双臂,阻碍了再组合,限制可扩展性。为此,我们提出SkillVLA,一种专为支持技能复用而设计的框架。大量实验表明,SkillVLA显著提升了技能组合能力,整体成功率从0%提升至51%,并在协作与长时序任务上表现优异。

原文摘要 · Abstract (English)

Recent progress in vision-language-action (VLA) models has demonstrated strong potential for dual-arm manipulation, enabling complex behaviors and generalization to unseen environments. However, mainstream bimanual VLA formulations largely overlook the critical challenge of combinatorial diversity. Different pairings of single-arm behaviors can induce qualitatively distinct task behaviors, yet existing models do not explicitly account for this structure. We argue that effective bimanual VLAs should support skill reuse - the ability to recombine previously learned single-arm skills across novel left-right pairings - thereby avoiding the need to separately learn every possible combination. Current VLA designs entangle skills across arms, preventing such recomposition and limiting scalability. To address this limitation, we propose SkillVLA, a framework explicitly designed to enable skill reuse in dual-arm manipulation. Extensive experiments demonstrate that SkillVLA substantially improves skill composition, increasing overall success rate from 0% to 51%, and achieves strong performance on cooperative and long-horizon tasks.

双臂操作技能复用视觉语言动作组合多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。