让视频生成更真实:通过几何一致性奖励减少物体变形和背景错乱
GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation

- 用光流与深度-姿态预测分离刚性运动与动态物体,评估其几何一致性
- 在多个动态场景中显著降低几何失真,同时保持视觉质量
- 适用于含相机移动和物体运动的复杂场景,可适配各类生成模型
视频生成中的几何一致性仍是难题:基于网络数据训练的文本到视频扩散模型仅隐式建模几何,导致物体形变、纹理漂移及相机运动下的非刚性背景问题。现有方法或仅作为副产品提升一致性,或仅适用于静态场景,或需完全重校准模型隐空间。本文提出一种几何一致性奖励,直接衡量生成视频中运动是否符合一致场景。核心思路是:物理一致视频中,背景运动应由刚性相机运动解释,独立移动物体应沿轨迹保持外观一致性。通过光流、深度-姿态预测与特征对应关系,分离刚性与动态区域并分别评估一致性。将该奖励结合强化学习微调,使几何一致性从涌现特性变为显式优化目标。方法对模型无依赖,适用于包含相机与物体运动的多样化动态场景。实验表明,在强基线基础上显著减少时间维度几何伪影,同时保持感知质量。代码与模型权重已公开。
原文摘要 · Abstract (English)
Generating geometrically consistent videos remains an open challenge: text-to-video diffusion models trained on web-scale data treat geometry only implicitly, leading to object deformation, texture drift, and non-rigid backgrounds under camera motion. Existing solutions either improve consistency as a byproduct, apply only to static scenes or realign the latent space of the model completely. We introduce a geometry-consistency reward that directly measures whether motion in a generated video is compatible with a coherent scene. Our key insight is that in physically consistent videos, background motion should be explainable by rigid camera-induced flow, while independently moving objects should preserve appearance identity along motion trajectories. We operationalize this using optical flow, depth--pose predictions, and feature-based correspondence to separate rigid and dynamic regions and evaluate their respective consistency. Integrating this reward with reinforcement fine-tuning transforms geometric consistency from an emergent property into an explicit optimization objective for video generators. The approach is model agnostic and applies to diverse dynamic scenes containing both camera and object motion. Experiments show substantial reductions in temporal geometric artifacts over strong baselines while preserving perceptual quality. Code and model weights are published.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。