将掩码图像建模与知识蒸馏引入双曲空间,高效捕捉视觉语义层级结构。
HMID-Net: An Exploration of Masked Image Modeling and Knowledge Distillation in Hyperbolic Space
- 在双曲空间中融合掩码图像建模与知识蒸馏,提升模型训练效率。
- 在图像分类与检索任务上显著超越MERU和CLIP等现有模型。
- 设计专用蒸馏损失函数,实现双曲空间中的有效知识迁移。
视觉与语义概念通常呈层次结构,例如'猫'这一概念涵盖所有猫的图像。近期研究MERU成功将多模态学习技术从欧几里得空间迁移到双曲空间,有效捕捉了视觉-语义层次结构。然而,如何更高效地训练模型以捕获并利用这种层次结构仍是关键问题。本文提出双曲掩码图像与蒸馏网络(HMID-Net),首次将掩码图像建模(MIM)与知识蒸馏技术结合于双曲空间,构建高效模型。我们设计了一种专用于双曲空间的知识蒸馏损失函数,促进有效知识传递。实验表明,双曲空间中的MIM与知识蒸馏可达到与欧几里得空间相当的效果。广泛评估显示,该方法在图像分类与检索等下游任务中显著优于MERU和CLIP等现有模型。
原文摘要 · Abstract (English)
Visual and semantic concepts are often structured in a hierarchical manner. For instance, textual concept `cat' entails all images of cats. A recent study, MERU, successfully adapts multimodal learning techniques from Euclidean space to hyperbolic space, effectively capturing the visual-semantic hierarchy. However, a critical question remains: how can we more efficiently train a model to capture and leverage this hierarchy? In this paper, we propose the Hyperbolic Masked Image and Distillation Network (HMID-Net), a novel and efficient method that integrates Masked Image Modeling (MIM) and knowledge distillation techniques within hyperbolic space. To the best of our knowledge, this is the first approach to leverage MIM and knowledge distillation in hyperbolic space to train highly efficient models. In addition, we introduce a distillation loss function specifically designed to facilitate effective knowledge transfer in hyperbolic space. Our experiments demonstrate that MIM and knowledge distillation techniques in hyperbolic space can achieve the same remarkable success as in Euclidean space. Extensive evaluations show that our method excels across a wide range of downstream tasks, significantly outperforming existing models like MERU and CLIP in both image classification and retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。