arXiv:2603.06698cs.CV2026-03

让小模型学大模型,用对比学习防止特征坍缩。

Breaking the Geometric Bottleneck: Contrastive Expansion in Asymmetric Cross-Modal Distillation

  • 引入对比损失扩展学生模型的特征流形,对抗知识蒸馏中的维度坍缩。
  • 在CIFAR-10上将有效维度从~16提升至~38,逼近小模型容量极限。
  • 揭示容量与密度的权衡:小模型反而是更鲁棒的语义滤波器。

在异构架构间进行知识蒸馏时常导致表征空间的严重几何约束。本文研究将全局视觉变换器(如CLIP和DINOv2)蒸馏到容量受限的CNN时出现的维度坍缩现象。通过严格中心化SVD与有效秩分析,首次在CIFAR-10上发现标准余弦蒸馏使表征坍缩至约16的有效秩。为逆转此现象,引入辅助对比目标(InfoNCE),使学生模型流形扩展2.4倍(达约38有效维度)。进一步表明,尽管DINOv2具有均匀几何结构部分抑制坍缩,但对比扩展仍是达到CNN拓扑容量极限(约82维)的普遍需求。最后揭示关键容量-密度权衡:固定流形内过度参数化导致脆弱性,而容量受限模型则充当最优低通语义滤波器,成功恢复固有噪声免疫能力。

原文摘要 · Abstract (English)

Knowledge distillation between asymmetric architectures often induces severe geometric constraints on the learned representation space. In this work, we investigate the Dimensional Collapse phenomenon when distilling global Vision Transformers (CLIP and DINOv2) into capacity-constrained CNNs. By employing strictly centered SVD and Effective Rank, we first demonstrate a capacity-agnostic phase transition on CIFAR-10 where standard cosine distillation collapses representations to an intrinsic Effective Rank of ~16. To reverse this, we integrate an auxiliary contrastive objective (InfoNCE), expanding the student's manifold by 2.4x (to ~38 effective dimensions). We further demonstrate that while DINOv2's uniform geometry partially prevents collapse, contrastive expansion remains a universal requirement to reach the CNN's topological capacity limit (~82 dimensions). Finally, we reveal a critical capacity-density trade-off: overparameterization within fixed manifolds induces brittleness, while capacity-constrained models act as optimal low-pass semantic filters, successfully recovering inherent noise immunity.

知识蒸馏对比学习特征坍缩小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。