arXiv:2503.13171cs.ROcs.AI2025-03被引 2

用视觉语言模型与混合规划生成海量机器人模仿学习数据

HybridGen: VLM-Guided Hybrid Planning for Scalable Data Generation of Imitation Learning

  • 先用VLM解析专家演示,拆分为精准控制和可规划部分
  • 生成数据量大,7个任务平均比现有方法高5%成功率
  • 无需特定数据格式,适合多种模仿学习算法

大规模且多样化的示范数据对提升机器人模仿学习的泛化能力至关重要。然而,在真实场景中生成复杂操作的此类数据极具挑战。本文提出HybridGen,一种融合视觉语言模型(VLM)与混合规划的自动化框架。该框架采用两阶段流程:首先,利用VLM解析专家演示,将任务分解为依赖专家的(以物体为中心的姿态变换,实现精确控制)和可规划段(通过路径规划生成多样化轨迹);其次,姿态变换显著扩充第一阶段数据。关键优势在于,HybridGen无需特定数据格式即可生成大量训练数据,具有广泛适用性,这一点在多个模仿学习算法上得到实证。在七个任务及其变体上的评估表明,使用HybridGen训练的智能体在性能与泛化能力上均有显著提升,平均优于当前最优方法5%。尤其在最困难的任务变体中,其平均成功率达59.7%,显著高于Mimicgen的49.5%。结果充分证明了该方法的有效性与实用性。

原文摘要 · Abstract (English)

The acquisition of large-scale and diverse demonstration data are essential for improving robotic imitation learning generalization. However, generating such data for complex manipulations is challenging in real-world settings. We introduce HybridGen, an automated framework that integrates Vision-Language Model (VLM) and hybrid planning. HybridGen uses a two-stage pipeline: first, VLM to parse expert demonstrations, decomposing tasks into expert-dependent (object-centric pose transformations for precise control) and plannable segments (synthesizing diverse trajectories via path planning); second, pose transformations substantially expand the first-stage data. Crucially, HybridGen generates a large volume of training data without requiring specific data formats, making it broadly applicable to a wide range of imitation learning algorithms, a characteristic which we also demonstrate empirically across multiple algorithms. Evaluations across seven tasks and their variants demonstrate that agents trained with HybridGen achieve substantial performance and generalization gains, averaging a 5% improvement over state-of-the-art methods. Notably, in the most challenging task variants, HybridGen achieves significant improvement, reaching a 59.7% average success rate, significantly outperforming Mimicgen's 49.5%. These results demonstrating its effectiveness and practicality.

机器人学习数据生成视觉语言模型模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。