用强化学习提升视频生成的相机控制精度和真实感
Geo-Align: Video Generation Alignment via Metric Geometry Reward

- 通过度量3D估计器提取生成视频的精确相机轨迹
- 在真实视频上实现优于监督学习基线的可控性与画质
- 无需成对数据,结合真实条件与合成轨迹训练
近年来,相机控制视频生成取得显著进展。然而现有视频重渲染方法主要依赖合成数据的监督微调,而真实世界同步多视角视频数据极度稀缺。当前范式在处理分布外的真实视频时泛化能力有限,模型难以准确遵循物理尺度和相机轨迹。为此,我们提出Geo-Align,首个专为相机控制视频重渲染设计的强化学习框架。基于预训练模型,通过感知度量奖励机制优化模型,引入度量3D估计算法从生成视频中提取精确相机轨迹,显式惩罚旋转与平移偏差。此外,设计基于真实条件视频与合成数据导出目标相机轨迹的数据流水线策略,避免对成对数据的依赖。大量实验表明,Geo-Align在相机可控性与视觉保真度方面均持续优于现有监督学习基线,验证了方法的有效性。
原文摘要 · Abstract (English)
Camera-controlled video generation has achieved remarkable progress in recent years. However, existing video-to-video re-rendering methods primarily rely on Supervised Fine-Tuning using synthetic datasets. At present, there is an extreme scarcity of synchronized, multi-view real-world video data. Consequently, the prevailing paradigm often exhibits limited generalization when processing out-of-distribution real-world videos, with models struggling to accurately adhere to physical scales and camera trajectories. To bridge this gap, we propose Geo-Align, the first Reinforcement Learning framework specifically designed for camera-controlled video re-rendering. Built upon a pretrained model, we optimize the model through a scale-aware perceptual reward mechanism. Specifically, we introduce a metric 3D estimator to extract precise camera trajectories from generated videos, explicitly penalizing deviations in rotation and translation. Furthermore, we meticulously designed a data pipeline strategy based on real-world conditioning videos and target camera trajectories derived from synthetic data, eliminating the reliance on paired data. Extensive experiments demonstrate that Geo-Align consistently outperforms existing supervised learning baselines in both precise camera controllability and visual fidelity, indicating the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。