arXiv:2603.23272cs.CVcs.MM2026-03被引 1

用因果干预提升多模态图像融合的鲁棒性,避免虚假关联。

Multi-Modal Image Fusion via Intervention-Stable Feature Learning

  • 基于因果干预设计三种探测策略,识别真实跨模态依赖。
  • 在多个数据集上超越现有方法,高维任务性能提升显著。
  • 适合需要跨模态稳定融合的医疗、遥感等场景应用。

多模态图像融合旨在将不同模态的互补信息整合为统一表征。现有方法主要优化模态间的统计相关性,常捕获受数据集影响的虚假关联,在分布偏移下性能下降。本文提出一种受因果原理启发的干预框架,以识别鲁棒的跨模态依赖。借鉴Pearl因果层次理论,设计三种原则性干预策略:i)空间分离扰动的互补掩码,测试模态能否真正互补缺失信息;ii)相同区域的随机掩码,识别在部分可观测下仍具信息量的特征子集;iii)模态丢弃,评估各模态不可替代的贡献。基于这些干预,引入因果特征集成器(CFI),通过自适应不变门控学习识别并优先选择在多种扰动模式下保持重要性的特征,从而捕捉稳健的模态依赖而非虚假相关。大量实验表明,该方法在公开基准和下游高层视觉任务中均达到最优性能。

原文摘要 · Abstract (English)

Multi-modal image fusion integrates complementary information from different modalities into a unified representation. Current methods predominantly optimize statistical correlations between modalities, often capturing dataset-induced spurious associations that degrade under distribution shifts. In this paper, we propose an intervention-based framework inspired by causal principles to identify robust cross-modal dependencies. Drawing insights from Pearl's causal hierarchy, we design three principled intervention strategies to probe different aspects of modal relationships: i) complementary masking with spatially disjoint perturbations tests whether modalities can genuinely compensate for each other's missing information, ii) random masking of identical regions identifies feature subsets that remain informative under partial observability, and iii) modality dropout evaluates the irreplaceable contribution of each modality. Based on these interventions, we introduce a Causal Feature Integrator (CFI) that learns to identify and prioritize intervention-stable features maintaining importance across different perturbation patterns through adaptive invariance gating, thereby capturing robust modal dependencies rather than spurious correlations. Extensive experiments demonstrate that our method achieves SOTA performance on both public benchmarks and downstream high-level vision tasks.

多模态融合因果推理图像融合鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。