arXiv:2505.10075cs.ROcs.CV2025-05被引 23

用显式运动流提升机器人操作的视觉世界模型预测能力

FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation

  • 以3D场景光流作为显式运动表示,分离预测与渲染过程
  • 在4个基准上实现7%语义相似度、11%像素质量、6%成功率提升
  • 适合需要精准视觉预测的机器人抓取与规划任务

本文研究用于机器人操作的更好视觉世界模型,即基于历史帧和机器人动作预测未来视觉观测。特别地,我们考虑基于RGB-D帧的世界模型(RGB-D世界模型)。与传统方法将动态预测隐式嵌入并统一于单个模型不同,我们提出FlowDreamer,采用3D场景光流作为显式运动表示。FlowDreamer首先通过U-Net从历史帧和动作条件中预测3D场景光流,再由扩散模型利用该光流生成未来帧。尽管结构模块化,但模型可端到端训练。我们在4个不同基准上进行实验,涵盖视频预测与视觉规划任务。结果表明,相比其他基线RGB-D世界模型,FlowDreamer在语义相似度上提升7%,像素质量提升11%,各类机器人操作场景下的任务成功率提升6%。

原文摘要 · Abstract (English)

This paper investigates training better visual world models for robot manipulation, i.e., models that can predict future visual observations by conditioning on past frames and robot actions. Specifically, we consider world models that operate on RGB-D frames (RGB-D world models). As opposed to canonical approaches that handle dynamics prediction mostly implicitly and reconcile it with visual rendering in a single model, we introduce FlowDreamer, which adopts 3D scene flow as explicit motion representations. FlowDreamer first predicts 3D scene flow from past frame and action conditions with a U-Net, and then a diffusion model will predict the future frame utilizing the scene flow. FlowDreamer is trained end-to-end despite its modularized nature. We conduct experiments on 4 different benchmarks, covering both video prediction and visual planning tasks. The results demonstrate that FlowDreamer achieves better performance compared to other baseline RGB-D world models by 7% on semantic similarity, 11% on pixel quality, and 6% on success rate in various robot manipulation domains.

世界模型机器人操作3D光流扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。