用可验证的几何奖励提升视频生成的摄像机控制精度
Taming Camera-Controlled Video Generation with Verifiable Geometry Reward
- 设计可验证几何奖励,通过分段对比相机轨迹实现细粒度反馈
- 在多场景下显著提升摄像机控制准确率与几何一致性
- 适合需要精确摄像机控制的视频生成研究者使用
近期视频扩散模型在摄像机控制视频生成方面取得显著进展,但多数方法仅依赖监督微调(SFT),在线强化学习(RL)后训练仍待探索。本文提出一种在线强化学习后训练框架,优化预训练视频生成器以实现精准摄像机控制。为提升强化学习有效性,设计可验证几何奖励:估计生成与参考视频的3D相机轨迹,将其分割为短段,计算每段相对位姿,通过对比生成-参考段对并赋予对齐得分作为奖励信号,缓解奖励稀疏性,提升优化效率。同时构建涵盖多样大振幅运动及动态主体的综合性数据集。大量实验表明,该方法在摄像机控制精度、几何一致性与视觉质量等方面均显著优于SFT基线,证明其在推动摄像机控制视频生成方面的优势。
原文摘要 · Abstract (English)
Recent advances in video diffusion models have remarkably improved camera-controlled video generation, but most methods rely solely on supervised fine-tuning (SFT), leaving online reinforcement learning (RL) post-training largely underexplored. In this work, we introduce an online RL post-training framework that optimizes a pretrained video generator for precise camera control. To make RL effective in this setting, we design a verifiable geometry reward that delivers dense segment-level feedback to guide model optimization. Specifically, we estimate the 3D camera trajectories for both generated and reference videos, divide each trajectory into short segments, and compute segment-wise relative poses. The reward function then compares each generated-reference segment pair and assigns an alignment score as the reward signal, which helps alleviate reward sparsity and improve optimization efficiency. Moreover, we construct a comprehensive dataset featuring diverse large-amplitude camera motions and scenes with varied subject dynamics. Extensive experiments show that our online RL post-training clearly outperforms SFT baselines across multiple aspects, including camera-control accuracy, geometric consistency, and visual quality, demonstrating its superiority in advancing camera-controlled video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。