arXiv:2511.16602cs.AI2025-11被引 3

用自省式训练框架,让智能体在少数据下高效学习复杂任务。

Bridging VLMs and Embodied Intelligence with Deliberate Practice Policy Optimization

  • 设计自省循环:交替进行监督微调与强化学习,自动识别弱点并精准投入资源。
  • 在有限数据下训练的模型比基线提升20.3%,超过100B参数开源模型10.6%。
  • 适合研究通用具身智能、数据稀缺场景下的高效训练方法的研究者。

构建通用且多功能的具身智能系统面临两大挑战:真实世界数据稀缺且成本高昂,以及现有方法算法效率低下,资源消耗过大。为此,我们提出自省式实践策略优化(DPPO),一种元认知“元环”训练框架,动态交替执行监督微调(能力扩展)和强化学习(技能精炼)。该框架可自动识别薄弱环节并实现资源的靶向分配,专为从稀疏、有限数据中最大化学习效率而设计。理论上,DPPO可形式化为统一的偏好学习框架。实验上,采用DPPO训练的视觉-语言具身模型Pelican-VL 1.0,在性能上相较基线模型提升20.3%,超越同规模(100B参数)的开源模型达10.6%。我们已开源模型与代码,提供首个系统性缓解数据与资源瓶颈的框架,助力社区高效构建多功能具身智能体。

原文摘要 · Abstract (English)

Developing a universal and versatile embodied intelligence system presents two primary challenges: the critical embodied data bottleneck, where real-world data is scarce and expensive, and the algorithmic inefficiency of existing methods, which are resource-prohibitive. To address these limitations, we introduce Deliberate Practice Policy Optimization (DPPO), a metacognitive ``Metaloop'' training framework that dynamically alternates between supervised fine-tuning (competence expansion) and reinforcement learning (skill refinement). This enables automatic weakness identification and targeted resource allocation, specifically designed to maximize learning efficiency from sparse, finite data. Theoretically, DPPO can be formalised as a unified preference-learning framework. Empirically, training a vision-language embodied model with DPPO, referred to as Pelican-VL 1.0, yields a 20.3% performance improvement over the base model and surpasses open-source models at the 100B-parameter scale by 10.6%. We are open-sourcing both the models and code, providing the first systematic framework that alleviates the data and resource bottleneck and enables the community to build versatile embodied agents efficiently.

具身智能自省训练高效学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。