arXiv:2510.26271cs.CL2025-10

小模型也能保持多语言能力,关键在蒸馏方法选择。

Distilling Multilingual Vision-Language Models: When Smaller Models Stay Multilingual

  • 对比五种蒸馏方法,找出维持多语言一致性的最优方案。
  • 模型减半后,部分方法仍能保持检索任务的多语言鲁棒性。
  • 适合关注多语言模型压缩与性能平衡的研究者。

视觉-语言模型(VLMs)在不同语言上的表现存在不均衡,模型缩小后问题更严重。尽管知识蒸馏(KD)在将大模型知识迁移至小模型方面表现良好,但其在多语言场景下的应用仍研究不足。本文对五种蒸馏方法进行受控实验,分析其对跨语言表征一致性及压缩后下游任务稳定性的影响。我们在CLIP和SigLIP2上测试了五种蒸馏配置,并在域内检索与域外视觉问答任务上评估效果。结果发现,某些配置在模型规模减半的情况下仍能保持或提升多语言检索鲁棒性,而另一些则无法维持跨任务稳定性,揭示出设计敏感的权衡关系,仅靠聚合准确率无法发现。

原文摘要 · Abstract (English)

Vision-language models (VLMs) exhibit uneven performance across languages, a problem that is often exacerbated when the model size is reduced. While Knowledge distillation (KD) demonstrates promising results in transferring knowledge from larger to smaller VLMs, applying KD in multilingualism is an underexplored area. This paper presents a controlled empirical study of KD behavior across five distillation approaches, isolating their effects on cross-lingual representation consistency and downstream performance stability under model compression. We study five distillation formulations across CLIP and SigLIP2, and evaluate them on in-domain retrieval and out-of-domain visual QA. We find that some configurations preserve or even improve multilingual retrieval robustness despite halving model size, but others fail to maintain cross-task stability, exposing design-sensitive trade-offs that aggregate accuracy alone does not reveal.

多语言模型压缩知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。