arXiv:2602.12636cs.LGcs.RO2026-02

用生成引导的对比奖励,让机器人高效学会复杂操作。

Dual-Granularity Contrastive Reward via Generated Episodic Guidance for Efficient Embodied RL

  • 基于视频模型生成任务指引,无需人工标注
  • 双粒度对比奖励提升探索效率和策略收敛性
  • 适合需要快速学习的真实世界机器人任务

在具身强化学习中,设计合适的奖励函数是关键挑战。轨迹成功奖励虽适用于人类评判或模型拟合,但稀疏性严重限制了样本效率。现有密集奖励方法依赖高质量人工标注数据或大量专家监督。为此,本文提出双粒度对比奖励框架DEG,通过大视频生成模型的先验知识,仅需少量专家视频即可为每轮RL生成专属任务指引。随后,提出的双粒度奖励在对比自监督潜在空间中,平衡粗粒度探索与细粒度匹配,引导智能体逐步逼近生成的指引视频,最终完成目标任务。在18个仿真与真实场景中的多样化任务上实验表明,DEG不仅能作为高效的探索刺激,帮助智能体快速发现稀疏的成功奖励,还能独立实现有效的强化学习与稳定的策略收敛。

原文摘要 · Abstract (English)

Designing suitable rewards poses a significant challenge in reinforcement learning (RL), especially for embodied manipulation. Trajectory success rewards are suitable for human judges or model fitting, but the sparsity severely limits RL sample efficiency. While recent methods have effectively improved RL via dense rewards, they rely heavily on high-quality human-annotated data or abundant expert supervision. To tackle these issues, this paper proposes Dual-granularity contrastive reward via generated Episodic Guidance (DEG), a novel framework to seek sample-efficient dense rewards without requiring human annotations or extensive supervision. Leveraging the prior knowledge of large video generation models, DEG only needs a small number of expert videos for domain adaptation to generate dedicated task guidance for each RL episode. Then, the proposed dual-granularity reward that balances coarse-grained exploration and fine-grained matching, will guide the agent to efficiently approximate the generated guidance video sequentially in the contrastive self-supervised latent space, and finally complete the target task. Extensive experiments on 18 diverse tasks across both simulation and real-world settings show that DEG can not only serve as an efficient exploration stimulus to help the agent quickly discover sparse success rewards, but also guide effective RL and stable policy convergence independently.

强化学习具身智能生成引导奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。