arXiv:2511.17886cs.CVcs.CL2025-11

强教师不等于好学生,CLIP知识蒸馏效果反降

When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA

  • 系统测试不同规模的CLIP教师模型进行知识蒸馏
  • 大模型教师反而导致视觉问答任务性能下降
  • 挑战了通用知识蒸馏有效性的主流认知

视觉语言模型(VLM)在多模态任务中取得显著进展,但其巨大的计算开销限制了高效部署。知识蒸馏(KD)已成为构建轻量级高性能模型的有效方法,在自然语言和视觉领域均有充分证据支持。然而,其在VLM,特别是CLIP类模型中的应用仍受限,通常仅限于小型教师模型和分类、检索等窄范围任务。本文首次系统研究了多种规模的CLIP类教师模型的知识蒸馏,从标准基线到大规模前沿模型。与自然语言处理和视觉领域的趋势相反,我们发现更强的教师并不一定产生更好的学生;现有蒸馏框架往往无法扩展,导致下游多模态任务(如视觉问答)性能下降。这一发现挑战了知识蒸馏领域的普遍假设,为设计参数高效的多模态模型指明了新方向。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have achieved remarkable success across multimodal tasks, yet their substantial computational demands hinder efficient deployment. Knowledge distillation (KD) has emerged as a powerful approach for building lightweight but competitive models, with strong evidence from both language and vision domains. However, its application to VLMs, particularly CLIP-style models, remains limited, often constrained to small-scale teachers and narrow evaluation tasks such as classification or retrieval. In this work, we present the first systematic study of distillation across a range of CLIP-style teacher models, ranging from standard baselines to large-scale state-of-the-art models. Contrary to trends observed in NLP and vision, we find that stronger teachers do not consistently yield better students; in fact, existing distillation frameworks often fail to scale, leading to degraded performance in downstream multimodal tasks such as visual question answering. Our findings challenge prevailing assumptions in KD and point toward new directions for designing parameter-efficient multimodal models.

知识蒸馏CLIP多模态视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。