arXiv:2608.12203cs.CV2026-08中稿 · ECCV

用几何先验加速自动驾驶视频生成,减少采样步骤

GeoFlow: Efficient Driving Video Generation via Geometry-Aligned Priors

论文配图:GeoFlow: Efficient Driving Video Generation via Geometry-Aligned Priors
图 1 · 摘自论文原文
  • 以多视角几何构建对齐的噪声初始分布,替代传统独立高斯噪声
  • 仅需数小时微调,即可实现少步生成高质量视频,推理步数大幅减少
  • 适合需要快速生成真实感驾驶视频的研究与应用者

生成模型如扩散模型和流匹配在合成高保真驾驶视频方面表现优异,但受制于高推理延迟,因需大量采样步骤。我们指出,这种低效源于对标准高斯源分布的依赖,即连续帧从独立高斯噪声初始化。该范式忽视了驾驶视频中固有的时空相关性,迫使模型从噪声中重复重建前帧存在的确定性场景结构,造成计算冗余且易产生几何不一致。为此,我们提出GeoFlow框架,通过显式利用几何先验实现高效驾驶视频生成。不采用标准高斯噪声采样,而是借助多视角几何与空间自适应噪声注入构建几何对齐先验(GAP)分布作为起点。此初始化弥合了源分布与数据分布间的差距,使采样轨迹更直接、更短。大量实验表明,GeoFlow在训练与推理效率上均有显著提升:仅需数小时微调基线模型即可显著提升少步生成质量,而完全收敛训练则极大减少达到顶尖视频生成效果所需的推理步数。

原文摘要 · Abstract (English)

Generative models like Diffusion Models and Flow Matching have demonstrated remarkable capabilities in synthesizing high-fidelity driving videos, but are severely constrained by high inference latency due to the requirement of extensive sampling steps. We argue that this inefficiency stems from the prevailing reliance on a standard Gaussian source distribution, where consecutive frames are initialized as independent Gaussian noise. This paradigm disregards the rich spatiotemporal correlations inherent in driving videos, compelling the model to regenerate deterministic scene structures existing in previous frames from noise, which is both computationally redundant and prone to geometric inconsistency. To address this problem, we propose GeoFlow, a novel framework designed to achieve efficient driving video generation by harnessing explicit geometric priors. Instead of sampling from standard Gaussian noise, we leverage multi-view geometry and spatially-adaptive noise injection to construct a Geometry-Aligned Prior (GAP) distribution as starting point. This initialization bridges the gap between source distribution and data distribution, yielding a significantly straighter and shorter sampling trajectory. Extensive experiments demonstrate that GeoFlow can achieve remarkable efficiency of both training and inference: merely several hours of fine-tuning on baseline models can significantly boost few-step generation quality, while fully converged training drastically reduces number of inference steps required for state-of-the-art video generation.

视频生成扩散模型几何先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。