arXiv:2603.26599cs.CV2026-03中稿 · ECCV被引 15

用4D隐空间奖励提升视频生成几何一致性,不改动预训练模型。

VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward

  • 在隐空间构建几何模型,直接从潜在表示解码场景结构。
  • 引入相机运动平滑与几何重投影一致双奖励,提升动态场景稳定性。
  • 无需重复VAE解码,效率高,适合真实世界复杂视频生成任务。

大规模视频扩散模型虽具出色视觉质量,但常缺乏几何一致性。已有方法通过增加模块或几何对齐改进,但前者损害模型泛化能力,后者局限于静态场景且依赖需反复解码的RGB空间奖励,计算开销大,难以适应高度动态的真实场景。为保留预训练能力同时提升几何一致性,我们提出VGGRPO(视觉几何GRPO),一种基于4D重建能力的隐空间几何引导后训练框架。该框架引入隐空间几何模型(LGM),将视频扩散潜变量与几何基础模型结合,实现从潜空间直接解码场景几何。在此基础上,采用两种互补奖励进行潜空间分组相对策略优化:相机运动平滑奖励抑制轨迹抖动,几何重投影一致性奖励确保多视角几何一致。在静态与动态基准测试中,VGGRPO显著提升相机稳定性、几何一致性与整体质量,且无需昂贵的VAE解码,证明了潜空间几何引导强化学习在世界一致视频生成中的高效性与灵活性。

原文摘要 · Abstract (English)

Large-scale video diffusion models achieve impressive visual quality, yet often fail to preserve geometric consistency. Prior approaches improve consistency either by augmenting the generator with additional modules or applying geometry-aware alignment. However, architectural modifications can compromise the generalization of internet-scale pretrained models, while existing alignment methods are limited to static scenes and rely on RGB-space rewards that require repeated VAE decoding, incurring substantial compute overhead and failing to generalize to highly dynamic real-world scenes. To preserve the pretrained capacity while improving geometric consistency, we propose VGGRPO (Visual Geometry GRPO), a latent geometry-guided framework for geometry-aware video post-training. VGGRPO introduces a Latent Geometry Model (LGM) that stitches video diffusion latents to geometry foundation models, enabling direct decoding of scene geometry from the latent space. By constructing LGM from a geometry model with 4D reconstruction capability, VGGRPO naturally extends to dynamic scenes, overcoming the static-scene limitations of prior methods. Building on this, we perform latent-space Group Relative Policy Optimization with two complementary rewards: a camera motion smoothness reward that penalizes jittery trajectories, and a geometry reprojection consistency reward that enforces cross-view geometric coherence. Experiments on both static and dynamic benchmarks show that VGGRPO improves camera stability, geometry consistency, and overall quality while eliminating costly VAE decoding, making latent-space geometry-guided reinforcement an efficient and flexible approach to world-consistent video generation.

视频生成扩散模型几何一致性强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。