用强化学习生成多样抓取数据,提升视觉-动作模型泛化能力
RL-Driven Data Generation for Robust Vision-Based Dexterous Grasping
- 用强化学习生成丰富接触的抓取轨迹,适配不同物体形状
- 在模拟数据上训练的视觉策略可泛化到未见过的物体
- 适合需要高效数据生成的机器人灵巧操作研究者
本文提出一种基于强化学习(RL)的数据增强方法,以提升视觉-动作(VA)模型在灵巧抓取任务中的泛化能力。尽管真实-仿真-真实框架通过少量真实演示生成大规模仿真数据已被证明有效,但在灵巧抓取场景中仍面临挑战:在多种物体形状下实现稳定多指接触困难。为此,我们利用强化学习生成覆盖多样化几何形状的高接触率抓取数据。遵循真实-仿真-真实的范式,抓取技能被建模为可参数化、可调节的参考轨迹,并通过强化学习学习残差策略进行优化。该模块化设计实现了轨迹级控制,既与真实示范一致,又能适应复杂物体几何。在仿真增强数据上训练的视觉条件策略展现出对未见物体的强大泛化性能,验证了该方法缓解VA模型训练数据瓶颈的潜力。
原文摘要 · Abstract (English)
This work presents reinforcement learning (RL)-driven data augmentation to improve the generalization of vision-action (VA) models for dexterous grasping. While real-to-sim-to-real frameworks, where a few real demonstrations seed large-scale simulated data, have proven effective for VA models, applying them to dexterous settings remains challenging: obtaining stable multi-finger contacts is nontrivial across diverse object shapes. To address this, we leverage RL to generate contact-rich grasping data across varied geometries. In line with the real-to-sim-to-real paradigm, the grasp skill is formulated as a parameterized and tunable reference trajectory refined by a residual policy learned via RL. This modular design enables trajectory-level control that is both consistent with real demonstrations and adaptable to diverse object geometries. A vision-conditioned policy trained on simulation-augmented data demonstrates strong generalization to unseen objects, highlighting the potential of our approach to alleviate the data bottleneck in training VA models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。