arXiv:2606.03201cs.CVcs.AI2026-06

用视频预测生成奖励,让智能体跨域学专家动作。

Reinforcement Learning from Cross-domain Videos with Video Prediction Model

论文配图:Reinforcement Learning from Cross-domain Videos with Video Prediction Model
图 1 · 摘自论文原文
  • 构建跨域视频预测模型,将观察映射到专家域
  • 在8个颜色域和3个体型域任务中超越基线
  • 适合模拟到真实机器人迁移的强化学习场景

跨视觉域的专家视频强化学习面临无奖励信号和域差距挑战。我们提出XIPER(跨域视频预测奖励),一种基于专家视频的奖励模型,适用于智能体外观因颜色、形态或仿真到真实差异而不同的情况。XIPER训练一个跨域视频预测模型,将智能体观测映射至专家域,并以预测似然作为奖励信号。在DMC Color Suite(8项任务)和DMC Body Suite(3项任务)上的实验表明,即使存在显著的域差距,XIPER仍持续优于基线。进一步在仿真到真实迁移数据集上的分析显示,仅使用模拟专家视频,XIPER即可为真实机器人观测生成有意义的奖励信号。代码、预训练模型、数据集及视频演示见项目主页:https://sites.google.com/view/xiper

原文摘要 · Abstract (English)

Reinforcement learning from expert videos across visually distinct domains is challenging due to the absence of reward signals and the presence of domain gaps. We introduce XIPER (Cross-domain Video Prediction Reward), a reward model for learning from expert videos collected in a visually different domain, where the agent's appearance differs due to factors such as color, morphology, or the sim-to-real gap. More specifically, XIPER trains a cross-domain video prediction model that maps agent observations into the expert domain and uses the prediction likelihood as a reward signal. Experiments on the DMC Color Suite (8 tasks) and DMC Body Suite (3 tasks) show that XIPER consistently outperforms baselines despite domain gaps such as differences in agent color and morphology. We further analyze XIPER on a sim-to-real transfer dataset, demonstrating that it produces meaningful reward signals for real-robot observations given only simulated expert videos. Code, pretrained models, datasets and video demonstrations can be found on our project webpage: https://sites.google.com/view/xiper

强化学习视频预测跨域迁移仿真到真实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。