通过物理解耦结构融合多模态图像,提升真实场景还原度。
SGPDFuse: Semantically-Guided Physics-Disentanglement General Multi-Modal Image Fusion

- 基于预训练模型构建语义-物理桥梁,分离场景不变属性与环境干扰。
- 在红外可见光、多焦点等6个基准上达最优性能,统一架构适用。
- 适合需要高保真图像融合的遥感、自动驾驶领域应用。
多模态图像融合(MMIF)旨在将互补传感器数据整合为单一表征,既保留场景本质真实性,又消除环境干扰。现有方法依赖盲特征聚合,虽能积累信号但难以区分核心内容与物理退化。本文提出SGPDFuse,通过基于预训练视觉基础模型的语义-物理参数桥(SPPB),利用内在-变化原理将输入映射至物理解耦的结构表示。为引导分解,引入语义对齐机制:在同基础模型特征空间中,通过余弦相似性显式锚定融合表征于显著语义特征,以保留关键目标;同时通过格拉姆矩阵正则化强制保持物理纹理保真度,严格消除不自然伪影。大量实验表明,SGPDFuse在红外-可见光、多焦点及多曝光等基准上均实现当前最优性能,且仅用单一架构。
原文摘要 · Abstract (English)
Multimodal image fusion (MMIF) aims to integrate complementary sensor data into a single representation that preserves intrinsic scene reality while eliminating environmental interferences. Most existing approaches rely on blind feature aggregation, which excels at signal accumulation but fails to distinguish essential content from physical degradations. We propose SGPDFuse, which bridges this gap by mapping inputs into a physics-disentangled structural representation via a Semantic-Physical Parametric Bridge (SPPB) built on pretrained vision foundation models, utilizing the Intrinsic-Variation principle to decouple invariant scene attributes from transient environmental factors. To guide this decomposition, we introduce a Semantic Alignment mechanism: we explicitly anchor the fused representation to salient semantic features in the same foundation model feature space via cosine similarity to preserve critical targets, while enforcing physical texture fidelity through Gram-matrix regularization to strictly eliminate unnatural artifacts. Extensive experiments demonstrate that SGPDFuse achieves state-of-the-art performance across infrared-visible, multi-focus, and multi-exposure benchmarks using a single architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。