arXiv:2503.23359cs.CV2025-03中稿 · CVPR被引 12

提出VideoFusion模型,实现多模态视频融合中的时空协同建模。

VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion

  • 设计差分增强模块实现跨模态信息交互
  • 在220段视频上达到更优的时序一致性表现
  • 适合做红外可见光视频融合的科研与工程人员

相较于图像,视频能更好反映真实世界采集过程并包含丰富的时间线索。然而,由于缺乏大规模多传感器视频数据集,现有研究主要聚焦多图融合,限制了视频融合的发展及统一建模时空依赖的难度。为此,我们构建了M3SVD基准数据集,包含220段时间同步、空间配准的红外-可见光视频,共153,797帧,填补数据空白。同时提出VideoFusion模型,通过跨模态互补性与动态时序特性,从多模态输入生成时空一致的视频。具体包括:1)差分增强模块促进跨模态信息交互与增强;2)完整的模态引导融合策略自适应整合多模态特征;3)双向时序协同注意力机制动态聚合前后向时序上下文,强化跨帧特征表示。实验表明,VideoFusion在序列上优于现有以图像为中心的融合范式,有效缓解了时序不一致与干扰问题。

原文摘要 · Abstract (English)

Compared to images, videos better reflect real-world acquisition and possess valuable temporal cues. However, existing multi-sensor fusion research predominantly integrates complementary context from multiple images rather than videos due to the scarcity of large-scale multi-sensor video datasets, limiting research in video fusion and the inherent difficulty of jointly modeling spatial and temporal dependencies in a unified framework. To this end, we construct M3SVD, a benchmark dataset with $220$ temporally synchronized and spatially registered infrared-visible videos comprising $153,797$ frames, bridging the data gap. Secondly, we propose VideoFusion, a multi-modal video fusion model that exploits cross-modal complementarity and temporal dynamics to generate spatio-temporally coherent videos from multi-modal inputs. Specifically, 1) a differential reinforcement module is developed for cross-modal information interaction and enhancement, 2) a complete modality-guided fusion strategy is employed to adaptively integrate multi-modal features, and 3) a bi-temporal co-attention mechanism is devised to dynamically aggregate forward-backward temporal contexts to reinforce cross-frame feature representations. Experiments reveal that VideoFusion outperforms existing image-oriented fusion paradigms in sequences, effectively mitigating temporal inconsistency and interference. Project and M3SVD: https://github.com/Linfeng-Tang/VideoFusion.

视频融合多模态时空建模红外可见光

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。