arXiv:2606.16856cs.RO2026-06被引 1

用视频模型生成伪标签,大幅减少人类标注需求。

Video-Based Optimal Transport for Feedback-Efficient Offline Preference-Based Reinforcement Learning

论文配图:Video-Based Optimal Transport for Feedback-Efficient Offline Preference-Based Reinforcement Learning
图 1 · 摘自论文原文
  • 基于视频基础模型的最优传输对齐视觉轨迹
  • 仅需少量标注即在多个任务上超越现有方法
  • 适合机器人等真实场景中低人工干预的强化学习

向强化学习代理传达复杂目标通常需要精细的奖励设计。基于人类反馈的强化学习(PbRL)提供了一种替代方案,但其可扩展性受高标注成本限制。受视频基础模型(ViFMs)进展启发,我们提出视频驱动的最优传输偏好(VOTP),一种半监督框架,仅需少量标注即可学习有效奖励函数。通过利用最优传输在ViFMs的丰富表征空间中对齐视觉轨迹,VOTP能为大量无标签数据生成高质量伪标签,显著降低人工监督需求。在运动和操作基准上的广泛实验表明,VOTP在有限反馈预算下优于当前最先进的离线PbRL方法。我们还验证了VOTP在视觉干扰下的鲁棒性,并在真实机器人任务中展示了其有效性,仅用极少人工输入即可学习有意义的奖励。

原文摘要 · Abstract (English)

Conveying complex objectives to reinforcement learning (RL) agents often requires meticulous reward engineering. Preference-based RL (PbRL) offers a promising alternative by learning reward functions from human feedback, but its scalability is hindered by high labeling costs. Inspired by advances in Video Foundation Models (ViFMs), we present Video-based Optimal Transport Preference (VOTP), a semi-supervised framework that learns effective reward functions from only a handful of labels. By leveraging optimal transport to align visual trajectories within the rich representation space of ViFMs, VOTP effectively generates high-fidelity pseudo-labels for large amounts of unlabeled data, substantially reducing human supervision. Extensive experiments across locomotion and manipulation benchmarks demonstrate the superiority of VOTP, which outperforms state-of-the-art offline PbRL methods under limited feedback budgets. We also showcase the robustness of VOTP in the presence of visual distractors and validate its utility on real robotic tasks, where it learns meaningful rewards with minimal human input.

强化学习视频生成半监督机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。