提出视频时序关联桥接模型,提升生成视频的连贯性。
Time-Correlated Video Bridge Matching
- 在扩散桥接框架中显式建模帧间时序依赖关系。
- 在插帧、图像转视频、视频超分任务上均超越基线方法。
- 适合需要高时序一致性的视频生成与编辑场景。
扩散模型在噪声到数据的生成任务中表现优异,能将高斯分布映射到复杂数据分布。然而,它们难以建模复杂分布间的变换,限制了在数据到数据任务中的应用。桥接匹配模型虽能解决分布间转换问题,但其在时序相关数据序列中的应用尚未探索,这对视频生成与操作任务至关重要,因保持时间连贯性尤为关键。为此,我们提出时间相关视频桥接匹配(TCVBM),将桥接匹配扩展至视频领域的时序数据序列。TCVBM在扩散桥接过程中显式建模序列内依赖关系,直接将时序相关性融入采样过程。我们在三个视频任务中对比了该方法与经典桥接匹配及扩散模型:帧插值、图像到视频生成、视频超分辨率。TCVBM在多个定量指标、基准数据集及人工评估中均取得更优表现。
原文摘要 · Abstract (English)
Diffusion models excel in noise-to-data generation tasks, providing a mapping from a Gaussian distribution to a more complex data distribution. However, they struggle to model translations between complex distributions, limiting their effectiveness in data-to-data tasks. While Bridge Matching models address this by finding the translation between data distributions, their application to time-correlated data sequences remains unexplored. This is a critical limitation for video generation and manipulation tasks, where maintaining temporal coherence is particularly important. To address this gap, we propose Time-Correlated Video Bridge Matching (TCVBM), a framework that extends Bridge Matching to time-correlated data sequences in the video domain. TCVBM explicitly models inter-sequence dependencies within the diffusion bridge, directly incorporating temporal correlations into the sampling process. We compare our approach to classical methods based on bridge matching and diffusion models for three video-related tasks: frame interpolation, image-to-video generation, and video super-resolution. TCVBM achieves superior performance across multiple quantitative metrics, benchmark datasets and human evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。