arXiv:2409.10080cs.CVcs.AI2024-09被引 6

提出自适应判别自编码器,实现清晰自然的多模态图像融合。

DAE-Fuse: An Adaptive Discriminative Autoencoder for Multi-Modality Image Fusion

  • 分两阶段设计判别自编码框架,抑制模态偏差
  • 在红外与可见光融合中保持细节清晰度,提升视觉质量
  • 首次拓展至视频域,保证帧间时序一致性

在夜间或低能见度等极端场景下,可靠感知对自动驾驶、机器人和监控至关重要。多模态图像融合,尤其是红外成像的融合,通过整合不同模态的互补信息,可显著提升场景理解与决策能力。然而,现有方法存在明显局限:基于GAN的方法常生成模糊图像,缺乏细粒度细节;基于AE的方法可能偏向特定模态,导致融合结果不自然。为此,我们提出DAE-Fuse,一种新颖的两阶段判别自编码器框架,可生成清晰且自然的融合图像。此外,我们首次将图像融合技术拓展至视频领域,同时保持帧间时序一致性,进一步提升自主导航所需的感知能力。在多个公开数据集上的大量实验表明,DAE-Fuse在多项基准测试中达到顶尖性能,并展现出在医学图像融合等任务中的优异泛化能力。

原文摘要 · Abstract (English)

In extreme scenarios such as nighttime or low-visibility environments, achieving reliable perception is critical for applications like autonomous driving, robotics, and surveillance. Multi-modality image fusion, particularly integrating infrared imaging, offers a robust solution by combining complementary information from different modalities to enhance scene understanding and decision-making. However, current methods face significant limitations: GAN-based approaches often produce blurry images that lack fine-grained details, while AE-based methods may introduce bias toward specific modalities, leading to unnatural fusion results. To address these challenges, we propose DAE-Fuse, a novel two-phase discriminative autoencoder framework that generates sharp and natural fused images. Furthermore, We pioneer the extension of image fusion techniques from static images to the video domain while preserving temporal consistency across frames, thus advancing the perceptual capabilities required for autonomous navigation. Extensive experiments on public datasets demonstrate that DAE-Fuse achieves state-of-the-art performance on multiple benchmarks, with superior generalizability to tasks like medical image fusion.

图像融合多模态自编码器视频处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。