用少对齐实现强跨模态融合,解决高分辨率遥感图像训练难题
Better with Less: Tackling Heterogeneous Multi-Modal Image Joint Pretraining via Conditioned and Degraded Masked Autoencoder

- 通过条件对比学习与降级重建,弱化异构模态间非对应特征干扰
- 在100万样本上预训练,性能超越依赖更大数据量的基础模型
- 适合做高分辨率光学与雷达图像联合建模的研究者和工程应用
跨异构模态的鲁棒表征学习仍是多模态视觉的核心挑战。以高分辨率光学与合成孔径雷达(SAR)联合预训练为例,其潜力受制于“异质-分辨率悖论”:精细空间尺度加剧了复杂雷达几何与非同源光学纹理之间的物理差异。现有中等分辨率下的刚性对齐方法迁移到高分辨率场景时,要么导致特征抑制以强制等价,要么因极端认知不确定性引发特征污染,均造成表征退化与负迁移。为此,本文提出CoDe-MAE,开创“少对齐、强协同”新范式。首先,光学锚定知识蒸馏(OKD)将SAR斑点噪声映射至纯语义流形,隐式正则化;其次,条件对比学习(CCL)利用梯度缓冲机制对齐共享共识,安全保留物理差异特征;同时,跨模态降级重建(CDR)主动剥离非同源光谱伪特征,截断固有病态映射,捕捉真实结构不变量。大量实验验证理论假设。在100万样本上预训练的CoDe-MAE展现卓越数据效率,有效防止表征退化,在多样单模态与双模态下游任务中达到新最佳性能,显著优于在海量数据上训练的基础模型。
原文摘要 · Abstract (English)
Learning robust representations across extremely heterogeneous modalities remains a fundamental challenge in multi-modal vision. As a critical and profound instantiation of this challenge, high-resolution (HR) joint optical and synthetic aperture radar (SAR) pretraining seeks modality synergy to mutually enhance single-source representations; its potential is severely hindered by the Heterogeneity-Resolution Paradox: finer spatial scales drastically amplify the physical divergence between complex radar geometries and non-homologous optical textures. Consequently, migrating medium-resolution-oriented rigid alignment paradigms to HR scenarios triggers either severe feature suppression to force equivalence, or feature contamination driven by extreme epistemic uncertainty. Both extremes inevitably culminate in profound representation degradation and negative transfer. To overcome this bottleneck, we propose CoDe-MAE, pioneering a \textit{better synergy with less alignment} philosophy. First, Optical-anchored Knowledge Distillation (OKD) implicitly regularizes SAR's speckle noise by mapping it into a pure semantic manifold. Building on this, Conditioned Contrastive Learning (CCL) utilizes a gradient buffering mechanism to align shared consensus while safely preserving divergent physical signatures. Concurrently, Cross-Modal Degraded Reconstruction (CDR) deliberately strips non-homologous spectral pseudo-features, truncating the inherently ill-posed mapping to capture true structural invariants. Extensive analyses validate our theoretical claims. Pretrained on 1M samples, CoDe-MAE demonstrates remarkable data efficiency, successfully preventing representation degradation and establishing new state-of-the-art performance across diverse single- and bi-modal downstream tasks, substantially outperforming foundation models scaled on vastly larger datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。