用实时验证提升机器人技能学习效率,减少人工调参。
Real-Time Verification of Embodied Reasoning for Generative Skill Acquisition
- 将数学推理验证思路迁移至机器人技能学习,动态生成任务提示和成功标准。
- 新方法使新任务成功率提升24%,已知任务提升36%,平均任务成功率提高21%。
- 适合需要高效自动化训练的具身智能研究者,尤其关注减少人工标注。
生成式技能学习使具身智能体能主动构建可扩展、持续演化的控制技能库,对大型决策模型的发展至关重要。以往方法常依赖通用智能体(如大语言模型)的监督信号,但在复杂3D环境中的有效性尚不明确;全面评估带来巨大计算开销,严重制约技能学习效率。受数学推理验证成功的启发,我们提出VERGSA(验证生成式技能学习中的具身推理),系统性地将实时验证原则融入具身技能学习。VERGSA实现:1)将数学推理验证机制延伸至具身学习,动态在提示中引入情境相关任务,并定义子任务与整体任务的成功度量;2)设计自动可扩展的奖励标注方案,通过迭代确定场景配置与子任务学习对整体技能获取的贡献,合成密集奖励信号。据我们所知,这是首个验证驱动的生成式技能学习综合性训练数据集,彻底消除繁琐的人工奖励工程。实验验证了该方法的有效性:1)示例任务池使平均任务成功率提升21%;2)验证模型使新任务成功率提升24%,已知任务提升36%;3)在验证质量上优于基于大语言模型的裁判基线。
原文摘要 · Abstract (English)
Generative skill acquisition enables embodied agents to actively learn a scalable and evolving repertoire of control skills, crucial for the advancement of large decision models. While prior approaches often rely on supervision signals from generalist agents (e.g., LLMs), their effectiveness in complex 3D environments remains unclear; exhaustive evaluation incurs substantial computational costs, significantly hindering the efficiency of skill learning. Inspired by recent successes in verification models for mathematical reasoning, we propose VERGSA (Verifying Embodied Reasoning in Generative Skill Acquisition), a framework that systematically integrates real-time verification principles into embodied skill learning. VERGSA establishes 1) a seamless extension from verification of mathematical reasoning into embodied learning by dynamically incorporating contextually relevant tasks into prompts and defining success metrics for both subtasks and overall tasks, and 2) an automated, scalable reward labeling scheme that synthesizes dense reward signals by iteratively finalizing the contribution of scene configuration and subtask learning to overall skill acquisition. To the best of our knowledge, this approach constitutes the first comprehensive training dataset for verification-driven generative skill acquisition, eliminating arduous manual reward engineering. Experiments validate the efficacy of our approach: 1) the exemplar task pool improves the average task success rates by 21%, 2) our verification model boosts success rates by 24% for novel tasks and 36% for encountered tasks, and 3) outperforms LLM-as-a-Judge baselines in verification quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。