让机器人双手协作更自然,统一感知与控制框架提升操作精度
CUBic: Coordinated Unified Bimanual Perception and Control Framework

- 将双臂协调建模为统一感知问题,通过共享表征自动实现独立与协作
- 在RoboTwin上任务成功率显著提升,协调精度优于现有方法
- 适合研究双臂机器人控制、视觉运动策略的学者与工程师
近期视觉-运动策略学习使机器人能直接从视觉输入进行控制。然而,将此类端到端学习从单臂扩展到双臂操作仍具挑战,因需兼顾独立感知与双臂协同交互。现有方法通常偏向单一策略——要么解耦双臂以避免干扰,要么强加跨臂耦合以保证协调,缺乏统一处理。我们提出CUBic:一种协调统一的双臂感知与控制框架,将双臂协调重新建模为统一的感知建模问题。CUBic学习一个共享的分词表示,连接感知与控制,其独立性与协调性由结构内生决定,而非人工设计的耦合。方法包含三个组件:单向感知聚合、通过两个共享映射码本的双向感知协调,以及统一的感知到控制扩散策略。在RoboTwin基准上的大量实验表明,CUBic持续超越标准基线,在协调精度和任务成功率上显著优于当前最优视觉运动基线。
原文摘要 · Abstract (English)
Recent advances in visuomotor policy learning have enabled robots to perform control directly from visual inputs. Yet, extending such end-to-end learning from single-arm to bimanual manipulation remains challenging due to the need for both independent perception and coordinated interaction between arms. Existing methods typically favor one side -- either decoupling the two arms to avoid interference or enforcing strong cross-arm coupling for coordination -- thus lacking a unified treatment. We propose CUBic, a Coordinated and Unified framework for Bimanual perception and control that reformulates bimanual coordination as a unified perceptual modeling problem. CUBic learns a shared tokenized representation bridging perception and control, where independence and coordination emerge intrinsically from structure rather than from hand-crafted coupling. Our approach integrates three components: unidirectional perception aggregation, bidirectional perception coordination through two codebooks with shared mapping, and a unified perception-to-control diffusion policy. Extensive experiments on the RoboTwin benchmark show that CUBic consistently surpasses standard baselines, achieving marked improvements in coordination accuracy and task success rates over state-of-the-art visuomotor baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。