解决红外可见光视频融合中的时间漂移问题,提升动态场景感知稳定性。
DRFusion: Drift-Resilient Temporally Consistent Infrared-Visible Video Fusion

- 将融合任务重构为历史条件下的运动生成,避免刚性对齐
- 在多个数据集上实现最佳融合质量与时间一致性,无明显漂移
- 适合需要稳定多模态视频感知的自动驾驶、安防场景
红外与可见光视频融合对于动态场景中的全面感知至关重要。然而,保持时间一致性仍面临巨大挑战。传统基于光流的方法常因几何刚性导致伪影;标准扩散模型通常逐帧处理,扩展至自回归设置时缺乏内在时间约束,易产生严重误差累积和漂移,微小伪影随时间放大。为此,我们提出一种抗漂移的视频融合方法,将任务重构为历史条件下的运动生成。引入稳定历史引导与软时间锚定,将时间一致性转化为谱滤波,隐式聚合运动动态而无需刚性对齐。同时,采用解耦结构-运动适应策略,通过两阶段训练与潜在空间精炼,衔接预训练先验与结构约束。大量实验表明,该方法在融合质量与时间稳定性上均达到当前最优水平。
原文摘要 · Abstract (English)
Infrared and visible video fusion is essential for achieving comprehensive perception in dynamic scenes. However, maintaining temporal consistency remains a formidable challenge. Conventional methods relying on optical flow often suffer from geometric rigidity and ghosting artifacts. Moreover, standard diffusion-based fusion models typically operate in a frame-by-frame manner; when extended to autoregressive settings, they lack intrinsic temporal constraints and are prone to severe error accumulation and drifting, where minor artifacts amplify over time. To address these limitations, we propose a drift-resilient video fusion method that reformulates the task as history-conditioned motion generation. We introduce Stabilized History Guidance and Soft Temporal Anchoring to reframe temporal consistency as spectral filtering, implicitly aggregating motion dynamics without rigid alignment. Furthermore, our Decoupled Structure-Motion Adaptation strategy bridges pre-trained priors and structural constraints via two-stage training and latent refinement. Extensive experiments demonstrate that our method achieves state-of-the-art performance in both fusion quality and temporal stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。