跨模态知识蒸馏新方法,解决弱语义关联下信息传递难题
Asymmetric Cross-Modal Knowledge Distillation: Bridging Modalities with Weak Semantic Consistency
- 提出非对称跨模态蒸馏框架,利用自监督学习动态匹配师生样本
- 在遥感场景分类任务中,6个模型架构上均超越7种现有方法
- 适合处理多源异构数据的跨模态学习,尤其适用于标注稀缺场景
跨模态知识蒸馏在强语义关联的配对模态上表现优异,但现实场景中配对数据稀少。为此,本文提出弱语义一致性下的非对称跨模态知识蒸馏(ACKD),旨在连接语义重叠有限的模态。基于最优传输理论,我们验证了从强到弱语义一致性会加剧知识传输成本。为此,提出SemBridge框架,包含学生友好的匹配模块与语义感知对齐模块:前者通过自监督学习获取语义知识,并动态选择相关教师样本;后者采用拉格朗日优化寻找最优传输路径。为推动研究,构建了基于多光谱(MS)与异构RGB图像的基准数据集,用于遥感场景分类。大量实验表明,在6个不同数据集、7种模型架构上,该方法性能优于现有7种方法。
原文摘要 · Abstract (English)
Cross-modal Knowledge Distillation has demonstrated promising performance on paired modalities with strong semantic connections, referred to as Symmetric Cross-modal Knowledge Distillation (SCKD). However, implementing SCKD becomes exceedingly constrained in real-world scenarios due to the limited availability of paired modalities. To this end, we investigate a general and effective knowledge learning concept under weak semantic consistency, dubbed Asymmetric Cross-modal Knowledge Distillation (ACKD), aiming to bridge modalities with limited semantic overlap. Nevertheless, the shift from strong to weak semantic consistency improves flexibility but exacerbates challenges in knowledge transmission costs, which we rigorously verified based on optimal transport theory. To mitigate the issue, we further propose a framework, namely SemBridge, integrating a Student-Friendly Matching module and a Semantic-aware Knowledge Alignment module. The former leverages self-supervised learning to acquire semantic-based knowledge and provide personalized instruction for each student sample by dynamically selecting the relevant teacher samples. The latter seeks the optimal transport path by employing Lagrangian optimization. To facilitate the research, we curate a benchmark dataset derived from two modalities, namely Multi-Spectral (MS) and asymmetric RGB images, tailored for remote sensing scene classification. Comprehensive experiments exhibit that our framework achieves state-of-the-art performance compared with 7 existing approaches on 6 different model architectures across various datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。