arXiv:2512.09297cs.RO2025-12中稿 · RSS 2026被引 2

仅用一次真实演示,就能生成数千个物理可行的双手操作数据。

One-Shot Real-World Demonstration Synthesis for Scalable Bimanual Manipulation

  • 将任务拆解为不变协调块与可变调整,通过视觉对齐和轻量优化生成新动作。
  • 在6个双臂任务中,新策略对新物体姿态和形状泛化能力强,显著优于现有基线。
  • 支持少样本扩展和零样本跨机器人迁移,适合追求高效真实训练的团队。

学习灵巧的双手操作策略严重依赖大规模高质量示范数据,但当前方法存在固有权衡:遥操作能提供物理真实的数据,却极度耗时;仿真合成虽可高效扩展,却面临仿真到现实的差距。我们提出 BiDemoSyn 框架,仅需一个真实世界示例即可合成丰富的接触型、物理可行的双手操作示范。核心思想是将任务分解为不变的协调块与依赖物体的可变调整,再通过视觉引导对齐和轻量轨迹优化进行适应。该方法可在数小时内生成数千个多样化且物理可行的示范,无需重复遥操作或依赖不完美的仿真。在六个双臂任务中,基于 BiDemoSyn 数据训练的策略对新物体姿态和形状表现出强泛化能力,显著优于近期强基线。此外,框架可自然扩展至少样本合成,提升物体层面多样性与分布外泛化能力,同时保持高数据效率。更重要的是,使用该数据训练的策略可实现零样本跨平台迁移,得益于以物体为中心的观测和简化的6-DoF末端执行器动作表示,使策略摆脱具体机械结构动态影响。BiDemoSyn 有效弥合了效率与真实性的鸿沟,为复杂双手操作的实用模仿学习提供了一条无需牺牲物理真实性的可扩展路径。

原文摘要 · Abstract (English)

Learning dexterous bimanual manipulation policies critically depends on large-scale, high-quality demonstrations, yet current paradigms face inherent trade-offs: teleoperation provides physically grounded data but is prohibitively labor-intensive, while simulation-based synthesis scales efficiently but suffers from sim-to-real gaps. We present BiDemoSyn, a framework that synthesizes contact-rich, physically feasible bimanual demonstrations from a single real-world example. The key idea is to decompose tasks into invariant coordination blocks and variable, object-dependent adjustments, then adapt them through vision-guided alignment and lightweight trajectory optimization. This enables the generation of thousands of diverse and feasible demonstrations within several hours, without repeated teleoperation or reliance on imperfect simulation. Across six dual-arm tasks, we show that policies trained on BiDemoSyn data generalize robustly to novel object poses and shapes, significantly outperforming recent strong baselines. Beyond the one-shot setting, BiDemoSyn naturally extends to few-shot-based synthesis, improving object-level diversity and out-of-distribution generalization while maintaining strong data efficiency. Moreover, policies trained on BiDemoSyn data exhibit zero-shot cross-embodiment transfer to new robotic platforms, enabled by object-centric observations and a simplified 6-DoF end-effector action representation that decouples policies from embodiment-specific dynamics. By bridging the gap between efficiency and real-world fidelity, BiDemoSyn provides a scalable path toward practical imitation learning for complex bimanual manipulation without compromising physical grounding.

双手操作模仿学习数据生成零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。