让双臂机器人更懂配合,提升复杂操作成功率与稳定性。
Co-VLA: Coordination-Aware Structured Action Modeling for Dual-Arm Vision-Language-Action Systems

- 引入结构化动作专家,显式建模双臂协作关系
- 实测紧耦合任务成功率提升27%,真实场景表现翻倍
- 无需额外硬件,兼容现有控制流程,适合工业部署
视觉-语言-动作(VLA)模型在单臂和双臂机器人操作中表现出强大能力。以往方法依赖端到端学习,通过大型视觉-语言骨干网络实现连续动作预测,但仅靠隐式协调难以保证复杂双臂任务的可靠、可解释和稳定行为。本文提出Co-VLA,一种具备协调意识的双臂操作框架,在VLA模型中引入显式结构先验。我们在先进视觉-语言骨干基础上,将单一动作头替换为专为双臂协调设计的结构化动作专家(SAE)。通过模块化协调感知损失,显式构建共享潜在变量(编码任务级协调意图)与残差潜在变量(捕捉每臂执行调整)。部署时,潜在感知控制器(LAC)实时解析表示,动态调节同步强度、执行不对称性、平滑度与安全约束,运行于关节指令层,兼容标准控制流程,无需力控或阻抗控制。仿真与真实世界基准测试表明,Co-VLA显著优于单体基线:紧耦合任务成功率提升27%,跨域真实场景性能从13%提升至27%以上,任务完成时间最多缩短25%。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models show strong capabilities in single and dual-arm robotic manipulation. Prior works show coordinated bimanual behaviors can emerge from end-to-end learning, leveraging large vision-language backbones with continuous action prediction. However, as bimanual tasks become tightly coupled and execution constraints become critical, implicit coordination alone is insufficient to ensure reliable, interpretable, and stable behavior. In this work, we propose Co-VLA, a coordination-aware bimanual manipulation framework introducing explicit structural priors into VLA models. We instantiate our method on a state-of-the-art vision-language backbone by replacing its monolithic action head with a Structured Action Expert (SAE) designed for bimanual coordination. Specifically, we introduce explicit structure at the action generation level with a modular coordination-aware loss that shapes shared and residual latents according to task-specific structures. The shared latent encodes task-level coordination intent, while residual latents capture execution adjustments for each arm. At deployment, a Latent-Aware Controller (LAC) interprets the learned representations to modulate synchronization strength, execution asymmetry, smoothness, and safety constraints in real time. LAC operates at the joint-command level and remains compatible with standard control pipelines without requiring force or impedance control. Experiments across simulation and real-world benchmarks show Co-VLA significantly outperforms monolithic baselines, achieving a 27% success rate gain in tight-coordination tasks, more than doubling performance in OOD real-world scenarios (from 13% to 27%), and reducing task completion time by up to 25%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。