用生成的物体中心流形设计奖励,提升视觉强化学习的泛化能力。
GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning
- 基于跨实体数据集生成物体中心流,提取低维特征
- 在10个仿真与真实任务中表现优于基线方法
- 适合需要少样本、强泛化能力的机器人操控场景
近期研究显示,视频生成模型可通过逆动力学推导有效机器人动作。然而,这些方法严重依赖生成数据质量,且因缺乏环境反馈,在精细操作上表现不佳。尽管基于视频的强化学习提升了策略鲁棒性,仍受限于视频生成的不确定性,以及训练扩散模型所需的大规模机器人数据集。为此,我们提出GenFlowRL,从多样化跨实体数据集中训练生成的流形中提取形状化的奖励信号。该方法利用低维、物体中心特征,从多样演示中学习可泛化且鲁棒的策略。在10个操作任务中,包括仿真与真实世界跨实体评估,结果表明GenFlowRL能有效利用生成物体中心流提取的操作特征,在多种复杂场景中持续实现更优性能。
原文摘要 · Abstract (English)
Recent advances have shown that video generation models can enhance robot learning by deriving effective robot actions through inverse dynamics. However, these methods heavily depend on the quality of generated data and struggle with fine-grained manipulation due to the lack of environment feedback. While video-based reinforcement learning improves policy robustness, it remains constrained by the uncertainty of video generation and the challenges of collecting large-scale robot datasets for training diffusion models. To address these limitations, we propose GenFlowRL, which derives shaped rewards from generated flow trained from diverse cross-embodiment datasets. This enables learning generalizable and robust policies from diverse demonstrations using low-dimensional, object-centric features. Experiments on 10 manipulation tasks, both in simulation and real-world cross-embodiment evaluations, demonstrate that GenFlowRL effectively leverages manipulation features extracted from generated object-centric flow, consistently achieving superior performance across diverse and challenging scenarios. Our Project Page: https://colinyu1.github.io/genflowrl
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。