arXiv:2605.21710cs.RO2026-05

仅用一次演示生成多样且物理合理的双臂操作数据,提升鲁棒性。

PGDG: Physically Grounded Data Generation for Robust Bimanual Policy Learning from a Single Demonstration

论文配图:PGDG: Physically Grounded Data Generation for Robust Bimanual Policy Learning from a Single Demonstration
图 1 · 摘自论文原文
  • 通过物理驱动采样与智能筛选迭代生成高质量恢复行为数据。
  • 仿真中成功率从38%提至93%,真实世界从35%提至82%。
  • 适合需要少样本、高鲁棒性的双臂操控任务研究者使用。

接触丰富的双臂操作任务的行为克隆仍具挑战,因多样化示范成本高昂,且微小扰动即可能导致系统进入无恢复监督的异常状态。本文提出PGDG,一种零样本数据筛选框架,可将单次示范扩展为紧凑、物理合理、成功且多样的恢复行为数据集,无需额外人工标注。PGDG在物理驱动采样器与数据集筛选器间迭代:筛选器选择信息量大、不冗余且可恢复的行为,更新采样分布以覆盖未充分探索的恢复模式;采样器则基于新分布生成物理合理的轨迹,并保留成功路径。为进一步提升数据质量,采用短时程采样控制对高风险状态重标矫正动作。在四个双臂操作任务中,PGDG在仿真和零样本真实迁移中均优于纯空间增强方法。在RotateBox-Pitch任务中,仿真成功率由38%提升至93%,真实世界从35%提升至82%。该方法还有效支持基础模型如GR00T的微调,成功率从46%升至77%。更多结果见官网:https://cunxid.github.io/PGDG/

原文摘要 · Abstract (English)

Behavior cloning for contact-rich bimanual manipulation remains challenging because diverse demonstrations are expensive to collect, and even small disturbances can push the system into off-manifold states where no recovery supervision is available. We propose PGDG, a data generation framework with zero-shot curation that expands a single demonstration into a compact dataset of physically plausible, successful, and diverse recovery behaviors without additional human labeling. PGDG iterates between a physics-grounded sampler and a dataset curator, where the curator selects informative, non-redundant, and recoverable behaviors to update the sampling distribution toward under-covered recovery modes, and the sampler draws physically plausible rollout candidates from this updated distribution and retains successful trajectories. To further improve data quality, PGDG applies short-horizon sampling-based control to relabel selected risky states with corrective actions. Across four bimanual manipulation tasks, PGDG consistently outperforms spatial-only augmentation in both simulation and zero-shot real-world transfer. On RotateBox-Pitch, success improves from 38% to 93% in simulation and from 35% to 82% in the real world. PGDG also enables effective foundation models fine-tuning such as GR00T, increasing success from 46% to 77%. Additional results are available in our website: https://cunxid.github.io/PGDG/.

双臂操控数据生成强化学习物理模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。