融合对比学习与知识蒸馏的掩码自编码,提升多模态表征能力
KDC-MAE: Knowledge Distilled Contrastive Mask Auto-Encoder
- 整合对比学习、知识蒸馏与掩码建模三类自监督范式
- 在多模态任务中实现优于单一方法的表征学习性能
- 适合需要高效多模态表示的学习者参考
本文尝试拓展自监督学习(SSL)范式,通过结合对比学习、自蒸馏(知识蒸馏)和掩码数据建模三种主流SSL框架,构建协同一致的联合表征学习机制。提出的新架构KDC-MAE,采用互补掩码策略以学习模块间对应关系,并设计加权融合方式协调不同目标。实验表明,对比掩码对应关系与知识蒸馏目标共同提升了多模态、多任务场景下的学习效果。
原文摘要 · Abstract (English)
In this work, we attempted to extend the thought and showcase a way forward for the Self-supervised Learning (SSL) learning paradigm by combining contrastive learning, self-distillation (knowledge distillation) and masked data modelling, the three major SSL frameworks, to learn a joint and coordinated representation. The proposed technique of SSL learns by the collaborative power of different learning objectives of SSL. Hence to jointly learn the different SSL objectives we proposed a new SSL architecture KDC-MAE, a complementary masking strategy to learn the modular correspondence, and a weighted way to combine them coordinately. Experimental results conclude that the contrastive masking correspondence along with the KD learning objective has lent a hand to performing better learning for multiple modalities over multiple tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。