arXiv:2506.03079cs.CV2025-06被引 19

用4D占据表示提升机器人视频生成质量与可控性

ORV: 4D Occupancy-centric Robot Video Generation

论文配图:ORV: 4D Occupancy-centric Robot Video Generation
图 1 · 摘自论文原文
  • 以4D语义占据为中间表示,融合动作先验与视觉先验
  • 相比最优方法,视频生成FVD降低18.8%,任务成功率提升6.4%
  • 支持多视角一致性生成,适合需要高保真仿真数据的研究

当前具身智能面临数据稀缺问题,传统模拟器缺乏视觉真实感。可控视频生成正成为有前景的数据引擎,但现有动作条件方法仍存在画质差、时序不一致、控制对齐弱及仅限单视角等问题。我们归因于稀疏控制输入与密集像素输出之间的表征鸿沟。为此提出ORV,一种以4D占据为中心的机器人视频生成框架,将动作先验与占据引导的视觉先验相耦合。具体地,通过动作专家自适应层归一化(AdaLN)将分段7-DoF动作与视频隐变量对齐,并注入2D占据渲染作为软引导。同时,针对具身场景缺乏占据数据的问题,我们构建了大规模高质量的4D语义占据数据集ORV-Data。在BridgeV2、DROID和RT-1上,ORV显著提升视频生成质量和可控性,实现18.8%更低的FVD、视觉规划成功率+3.5%、策略学习成功率+6.4%。超越单视角生成,ORV原生支持多视角一致性合成,并可在显著域差距下实现仿真到现实迁移。代码、模型与数据见:https://orangesodahub.github.io/ORV

原文摘要 · Abstract (English)

Recent embodied intelligence suffers from data scarcity, while conventional simulators lack visual realism. Controllable video generation is emerging as a promising data engine, yet current action-conditioned methods still fall short: generated videos are limited in fidelity and temporal consistency, poorly aligned with controls, and often constrained to singleview settings. We attribute these issues to the representational gap between sparse control inputs and dense pixel outputs. Thus, we introduce ORV, a 4D occupancy-centric framework for robot video generation that couples action priors with occupancy-derived visual priors. Concretely, we align chunked 7-DoF actions with video latents via an Action-Expert AdaLN modulation, and inject 2D renderings of 4D semantic occupancy into the generation process as soft guidance. Meanwhile, a central obstacle is the lack of occupancy data for embodied scenarios; we therefore curate ORV-Data, a large-scale, high-quality 4D semantic occupancy dataset of robot manipulation. Across BridgeV2, DROID, and RT-1, ORV improves video generation quality and controllability, achieving 18.8% lower FVD than state of the art, +3.5% success rate on visual planning, and +6.4% success rate on policy learning. Beyond singleview generation, ORV natively supports multiview consistent synthesis and enables simulation-to-real transfer despite significant domain gaps. Code, models, and data are at: https://orangesodahub.github.io/ORV

视频生成4D占据机器人仿真可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。