arXiv:2604.01900cs.CV2026-04

通过频域分解与时间扰动,提升红外可见光视频融合的稳定性和细节保留。

FTPFusion: Frequency-Aware Infrared and Visible Video Fusion with Temporal Perturbation

  • 分高频低频建模,稀疏跨模态交互捕捉运动细节。
  • 引入时间扰动策略,增强对闪烁、抖动等干扰的鲁棒性。
  • 适合智能监控与低光环境下的视频融合任务。

红外与可见光视频融合在智能监控和低光监测中至关重要。然而,在保持时空稳定性的同时保留空间细节仍是核心挑战。现有方法或仅关注帧级增强而缺乏时间建模,或依赖复杂的时空聚合,常牺牲高频细节。本文提出FTPFusion,一种基于时间扰动与稀疏跨模态交互的频域感知融合方法。该方法将特征表示分解为高低频成分进行协同建模:高频分支通过稀疏跨模态时空交互捕捉运动相关上下文与互补细节;低频分支引入时间扰动策略,提升对闪烁、抖动及局部错位等复杂视频变化的鲁棒性。此外,设计了偏移感知的时间一致性约束,显式稳定受时间扰动影响的跨帧表示。在多个公开基准上的大量实验表明,FTPFusion在空间保真度与时间一致性多项指标上均持续优于当前最优方法。源代码将发布于 https://github.com/ixilai/FTPFusion。

原文摘要 · Abstract (English)

Infrared and visible video fusion plays a critical role in intelligent surveillance and low-light monitoring. However, maintaining temporal stability while preserving spatial detail remains a fundamental challenge. Existing methods either focus on frame-wise enhancement with limited temporal modeling or rely on heavy spatio-temporal aggregation that often sacrifices high-frequency details. In this paper, we propose FTPFusion, a frequency-aware infrared and visible video fusion method based on temporal perturbation and sparse cross-modal interaction. Specifically, FTPFusion decomposes the feature representations into high-frequency and low-frequency components for collaborative modeling. The high-frequency branch performs sparse cross-modal spatio-temporal interaction to capture motion-related context and complementary details. The low-frequency branch introduces a temporal perturbation strategy to enhance robustness against complex video variations, such as flickering, jitter, and local misalignment. Furthermore, we design an offset-aware temporal consistency constraint to explicitly stabilize cross-frame representations under temporal disturbances. Extensive experiments on multiple public benchmarks demonstrate that FTPFusion consistently outperforms state-of-the-art methods across multiple metrics in both spatial fidelity and temporal consistency. The source code will be available at https://github.com/ixilai/FTPFusion.

视频融合红外可见光时序稳定频域建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。