arXiv:2606.13279cs.RO2026-06被引 1

通过分层分解视觉选择与双臂交互,提升机器人双手操作的适应性与成功率。

See Selectively, Act Adaptively: Dual-Level Structural Decomposition for Bimanual Robot Manipulation

论文配图:See Selectively, Act Adaptively: Dual-Level Structural Decomposition for Bimanual Robot Manipulation
图 1 · 摘自论文原文
  • 双级结构分解:动态选择视觉输入,分离协调与独立动作路径。
  • 仿真中成功率提升27.7%,真实场景提升43.3%,优于单模块模型。
  • 适合复杂双臂任务,尤其对视觉变化大、交互模式多变的场景有效。

在双臂机器人操作中,任务相关的视觉信息随阶段和上下文变化,双臂交互模式也在独立与协同间切换,导致策略学习困难。现有单一视觉-语言-动作(VLA)策略通过共享表征和统一动作生成路径处理多样输入与交互模式,常无法分别建模视觉相关性与双臂交互结构。为此,我们提出基于双级结构分解的双臂操作VLA框架。视图选择性视觉路由模块动态调整腕部视角贡献,突出相关视觉线索;交互感知动作混合专家(MoE)将动作生成分解为协同与单臂路径,以适应不同双臂交互模式。我们在RoboTwin 2.0的六个模拟任务及三个长时程真实世界任务上评估该方法。相比单体基线,仿真平均成功率提升27.7%,真实世界提升43.3%,且在两种设置下均持续优于单模块变体。结果表明,联合考虑选择性视觉处理与显式双臂交互结构分解,为鲁棒双臂操作提供了有效归纳偏置。

原文摘要 · Abstract (English)

In bimanual robotic manipulation, task-relevant visual information varies with the task stage and context, while the interaction of the two arms shifts between independent and coordinated modes, making policy learning challenging. However, existing monolithic Vision-Language-Action (VLA) policies process diverse visual inputs and interaction patterns through a single shared representation and action generation pathway, often failing to separately account for visual relevance and bimanual interaction structure. To address this issue, we propose a bimanual manipulation VLA framework based on Dual-Level Structural Decomposition. The View-Selective Visual Router dynamically adjusts wrist-view contributions to emphasize relevant visual cues, while the Interaction-Aware Action Mixture-of-Experts (MoE) decomposes action generation into coordinated and arm-wise pathways to adapt to varying bimanual interaction modes. We evaluate the proposed method on six simulated bimanual manipulation tasks in RoboTwin 2.0 and three long-horizon real-world tasks. Our model improves the overall average success rate over a monolithic baseline by 27.7% in simulation and 43.3% in real-world evaluation, while consistently outperforming single-module variants across both settings. These results demonstrate that jointly considering selective visual processing and explicit decomposition of bimanual interaction structures provides an effective inductive bias for robust bimanual manipulation.

双臂操作视觉路由动作混合专家机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。