用几何先验指导视频生成,让物体运动更符合3D结构。
VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation
- 通过几何基础模型自动生成偏好信号,引导视频扩散模型。
- 仅用少量偏好对就显著提升时空稳定性和几何合理性。
- 无需人工标注,适合追求真实感视频生成的研究者。
尽管近期视频扩散模型(VDMs)生成效果视觉惊艳,但其在保持3D结构一致性方面存在根本缺陷,常导致物体形变或空间漂移。我们假设这些失败源于标准去噪目标缺乏几何一致性的显式激励。为此,提出VideoGPA(Video Geometric Preference Alignment),一种数据高效的自监督框架,利用几何基础模型自动推导密集偏好信号,通过直接偏好优化(DPO)引导VDM。该方法有效将生成分布推向内在的3D一致性,无需人工标注。VideoGPA在大量实验中显著提升时序稳定性、几何合理性与运动连贯性,以极小偏好对数量持续超越当前最优基线。
原文摘要 · Abstract (English)
While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deformation or spatial drift. We hypothesize that these failures arise because standard denoising objectives lack explicit incentives for geometric coherence. To address this, we introduce VideoGPA (Video Geometric Preference Alignment), a data-efficient self-supervised framework that leverages a geometry foundation model to automatically derive dense preference signals that guide VDMs via Direct Preference Optimization (DPO). This approach effectively steers the generative distribution toward inherent 3D consistency without requiring human annotations. VideoGPA significantly enhances temporal stability, geometric plausibility, and motion coherence using minimal preference pairs, consistently outperforming state-of-the-art baselines in extensive experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。