用几何约束提升视频生成的稳定性与一致性
Epipolar Geometry Improves Video Generation Models
- 通过成对极几何约束优化扩散模型,直接纠正运动轨迹不稳问题
- 使极线误差降低31%,人工评估一致性从54%提升至72%
- 适合关注视频生成几何一致性的研究者与应用开发者
视频生成模型虽在潜在扩散变换器与修正流训练下取得显著进展,但仍面临几何不一致、运动不稳定和视觉伪影等问题,破坏真实三维场景的沉浸感。本文探索如何通过极几何约束改善现代视频扩散模型。尽管训练数据庞大,现有模型仍未能捕捉基本几何规律。我们利用基于偏好优化的成对极几何约束,对扩散模型进行对齐,以数学上严谨的方式强制执行几何原则,有效解决轨迹不稳定与几何伪影问题。该方法无需端到端可微,计算高效。实验表明,经典几何约束提供的优化信号比现代学习度量更稳定。在静态场景配合动态相机的训练下,模型在多种动态场景中实现良好泛化。通过融合数据驱动学习与经典计算机视觉,本方法将极线误差减少31%,人工评估一致性从54%提升至72%,且不牺牲视觉质量。
原文摘要 · Abstract (English)
Video generation models have advanced significantly through the latent diffusion transformers trained with rectified flow techniques. Yet these models still struggle with geometric inconsistencies, unstable motion, and visual artifacts that break the illusion of realistic 3D scenes. 3D-consistent video generation could significantly impact numerous downstream applications in generation and reconstruction tasks. We explore how epipolar geometry constraints improve modern video diffusion models. Despite using massive training data, these models fail to capture fundamental geometric principles. We align diffusion models using pairwise epipolar geometry constraints via preference-based optimization, directly addressing unstable trajectories and geometric artifacts through mathematically principled geometric enforcement. Our approach efficiently enforces geometric principles without requiring end-to-end differentiability. Evaluation demonstrates that classical geometric constraints provide more stable optimization signals than modern learned metrics. Training on static scenes with dynamic cameras ensures metric quality while the model generalizes to various dynamic scenes. By bridging data-driven learning with classical computer vision, we reduce epipolar error by 31% and improve human-rated consistency from 54% to 72% without compromising visual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。