arXiv:2605.16937cs.CV2026-05

提出首个在线策略梯度的极端视角视频生成方法,解决大视角运动下的生成难题。

DEVIS-GRPO: Unleashing GRPO on Dynamic Extreme View Synthesis

论文配图:DEVIS-GRPO: Unleashing GRPO on Dynamic Extreme View Synthesis
图 1 · 摘自论文原文
  • 通过累积小视角增量实现大视角运动,无需昂贵配对数据
  • 在Kubric-4D上提升21.57% PSNR、7.31% SSIM,iPhone上LPIPS降18.56%
  • 适合需要高自由度轨迹控制的视频生成研究者

轨迹控制的视频生成对可控视频生成至关重要。现有方法在小视角相机运动下表现良好,但在大视角运动下性能显著下降。传统极端视角合成通常需专用视频对,标注成本高。为此,我们提出基于GRPO的动态极端视角视频生成框架DEVIS-GRPO,是首个用于极端视角视频生成的在线策略梯度方法。核心是新型采样策略:累积式动态极端视角合成(ADEVIS),通过逐步累积小视角增量实现大视角运动。该方法带来两大优势:1)提升训练效率,无需收集昂贵的大视角配对视频来预热策略模型;2)增强采样多样性,通过灵活调整轨迹配置实现。此外,设计多级一致性-质量奖励函数,筛选高质量样本用于模型优化。在Kubric-4D、iPhone和DL3DV数据集上的实验表明,本方法具有明显优势。在Kubric-4D上,非遮挡区域相比次优方法,PSNR提升21.57%,SSIM提升7.31%;在iPhone数据集上,LPIPS降低18.56%。

原文摘要 · Abstract (English)

Trajectory-controlled video generation has become essential for controllable video generation. While current methods perform well under small-view camera motions, they degrade significantly with large-view motions. Existing solutions for extreme-view synthesis typically require dedicated video pairs, demanding substantial annotation effort. To address these limitations, we propose Dynamic Extreme VIew Synthesis-GRPO (DEVIS-GRPO), a GRPO-based framework for trajectory-controlled video generation, the first online policy gradient method for extreme view video generation. Central to our approach is a novel sampling strategy: Accumulative Dynamic Extreme VIew Synthesis (ADEVIS), which achieves large-view camera motions by progressively accumulating small-view increments. This method delivers two key advantages: 1) enhanced training efficiency, as it eliminates the need to warm-start the policy model by collecting expensive paired large-view videos, and 2) increased sampling diversity, achieved by flexibly varying trajectory configurations. Finally, we designed a multi-level consistency-quality reward function to select high-quality samples for model optimization. Experiments on the Kubric-4D, iPhone, and DL3DV datasets demonstrate our method's superiority. On Kubric-4D, we achieve relative improvements of 21.57% in PSNR and 7.31% in SSIM over the second-best method in non-occlusion areas. On iPhone, LPIPS is reduced by 18.56%.

视频生成轨迹控制极端视角强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。