arXiv:2507.19817cs.RO2025-07中稿 · IROS 2025, oral pr…

让机器人用双眼看懂双手协作,零样本完成复杂操作。

Ag2x2: Robust Agent-Agnostic Visual Representations for Zero-Shot Bimanual Manipulation

  • 构建协调感知的视觉表征,同时捕捉物体状态与手部运动模式。
  • 在13个任务上达73.5%成功率,优于有专家奖励的基线方法。
  • 无需人工示范或奖励设计,适合大规模无监督技能学习。

双手操作是人类日常活动的核心,但其协同控制的复杂性使其仍具挑战。近期研究通过人类视频提取无代理依赖的视觉表征,实现了单臂操作的零样本学习;然而,这些方法忽略了双手协调所需的代理特异性信息(如末端执行器位置)。本文提出Ag2x2,一种面向双手操作的计算框架,通过协调感知的视觉表征,联合编码物体状态与手部运动模式,同时保持无代理特性。大量实验表明,Ag2x2在Bi-DexHands和PerAct2数据集上的13项多样化双手任务中实现73.5%的成功率,包括绳索等柔体物体的复杂场景。该性能不仅超越基线方法,更超过使用人工设计奖励训练的策略。此外,证明了Ag2x2所学表征可有效用于模仿学习,建立无需专家监督的可扩展技能获取流程。在无需人类示范或工程化奖励的前提下,保持跨任务鲁棒表现,推动复杂双手机器人技能的可扩展学习。

原文摘要 · Abstract (English)

Bimanual manipulation, fundamental to human daily activities, remains a challenging task due to its inherent complexity of coordinated control. Recent advances have enabled zero-shot learning of single-arm manipulation skills through agent-agnostic visual representations derived from human videos; however, these methods overlook crucial agent-specific information necessary for bimanual coordination, such as end-effector positions. We propose Ag2x2, a computational framework for bimanual manipulation through coordination-aware visual representations that jointly encode object states and hand motion patterns while maintaining agent-agnosticism. Extensive experiments demonstrate that Ag2x2 achieves a 73.5% success rate across 13 diverse bimanual tasks from Bi-DexHands and PerAct2, including challenging scenarios with deformable objects like ropes. This performance outperforms baseline methods and even surpasses the success rate of policies trained with expert-engineered rewards. Furthermore, we show that representations learned through Ag2x2 can be effectively leveraged for imitation learning, establishing a scalable pipeline for skill acquisition without expert supervision. By maintaining robust performance across diverse tasks without human demonstrations or engineered rewards, Ag2x2 represents a step toward scalable learning of complex bimanual robotic skills.

双手操作零样本学习视觉表征机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。