arXiv:2606.22471cs.RO2026-06

用强化学习生成多样高质数据,提升双臂灵巧操作的泛化能力

Scalable Multi-Task Data Generation via Reinforcement Learning for Language-Conditioned Bimanual Dexterous Manipulation

论文配图:Scalable Multi-Task Data Generation via Reinforcement Learning for Language-Conditioned Bimanual Dexterous Manipulation
图 1 · 摘自论文原文
  • 基于通用奖励设计与领域随机化,实现可扩展的数据生成
  • 在3个任务上显著提升语言指令下多任务策略的泛化性能
  • 适合研究灵巧操作、多任务学习和机器人数据合成的研究者

训练通用双臂灵巧操作策略的关键瓶颈在于缺乏大规模高质量数据集。模拟中的合成数据生成可克服形态不匹配、物理交互缺失及机器人动作生成等问题,提供比人类视频示范更可扩展的替代方案。然而,现有基于人类遥操作的方法任务多样性有限,因以物体为中心的轨迹匹配常忽略机器人执行可行性。强化学习(RL)虽具更广可扩展性,但常受限于人工设计的任务特定奖励。本文提出一种系统化的基于RL的数据生成流程,整合可泛化奖励设计、有效领域随机化及语言条件任务标注,生成多样化高质量数据,支持语言条件多任务策略训练。实验表明,生成数据显著提升三个代表性操作任务上的泛化能力。

原文摘要 · Abstract (English)

A key bottleneck in training generalist policies for bimanual dexterous manipulation is the lack of large-scale, high-quality datasets. Synthetic data generation in simulation provides a scalable alternative to human video demonstrations by overcoming challenges such as morphology mismatch, missing physical interactions, and the generation of robot actions. However, existing approaches based on human teleoperation offer limited task diversity, as object-centric trajectory matching often neglects the feasibility of robot execution. Reinforcement learning (RL) enables broader scalability but is often constrained by handcrafted, task-specific rewards. In this work, we propose a systematic RL-based data generation pipeline that integrates generalizable reward design, effective domain randomization, and language-conditioned task annotations. This pipeline synthesizes diverse, high-quality datasets for dexterous bimanual manipulation and enables training of language-conditioned multi-task policies. Our experiments show that the generated data significantly improves generalization across three representative manipulation tasks.

灵巧操作强化学习多任务学习数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。