统一视频融合框架,解决动态画面闪烁与不一致问题
A Unified Solution to Video Fusion: From Multi-Frame Learning to Benchmarking
- 利用多帧学习和光流特征扭曲实现时序连贯融合
- 在四个任务上均达当前最优,有效减少画面闪烁
- 适合视频处理、医疗影像等需要稳定时序的场景
真实世界是动态的,但大多数图像融合方法独立处理静态帧,忽略视频中的时间相关性,导致画面闪烁和时序不一致。为此,我们提出统一视频融合(UniVF)框架,通过多帧学习和基于光流的特征扭曲,实现信息丰富且时序一致的视频融合。为支持其开发,我们还构建了首个综合性视频融合基准(VF-Bench),涵盖四类任务:多曝光、多焦点、红外-可见光及医学图像融合。VF-Bench 提供高质量、对齐良好的视频对,来自合成数据生成与现有数据集严格筛选,并采用统一评估协议,联合衡量空间质量与时间一致性。大量实验表明,UniVF 在所有任务上均达到当前最优性能。
原文摘要 · Abstract (English)
The real world is dynamic, yet most image fusion methods process static frames independently, ignoring temporal correlations in videos and leading to flickering and temporal inconsistency. To address this, we propose Unified Video Fusion (UniVF), a novel and unified framework for video fusion that leverages multi-frame learning and optical flow-based feature warping for informative, temporally coherent video fusion. To support its development, we also introduce Video Fusion Benchmark (VF-Bench), the first comprehensive benchmark covering four video fusion tasks: multi-exposure, multi-focus, infrared-visible, and medical fusion. VF-Bench provides high-quality, well-aligned video pairs obtained through synthetic data generation and rigorous curation from existing datasets, with a unified evaluation protocol that jointly assesses the spatial quality and temporal consistency of video fusion. Extensive experiments show that UniVF achieves state-of-the-art results across all tasks on VF-Bench. Project page: https://vfbench.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。