arXiv:2510.00515cs.CV2025-10NeurIPS被引 22

通过渐进式一致性蒸馏,高效压缩视觉令牌提升多模态大模型性能

Efficient Multi-modal Large Language Models via Progressive Consistency Distillation

  • 分步压缩视觉令牌,逐层引入教师模型引导训练
  • 在多个数据集上实现更高效率与更强泛化能力
  • 适合追求模型轻量化与推理速度的开发者

多模态大模型中的视觉令牌消耗大量计算资源,严重制约其效率。现有方法虽尝试通过训练阶段压缩视觉令牌来提升效率,或修改模型结构或引入额外参数,但常忽视压缩带来的学习难度增加——特征空间剧烈扰动使模型参数难以快速适应。本文提出一种渐进式一致性蒸馏(EPIC)框架,将令牌压缩引起的特征空间扰动分解为逐令牌和逐层维度,分别设计令牌一致性蒸馏与层一致性蒸馏,借助教师模型指导并沿渐进学习轨迹降低训练难度。大量实验表明,该框架在效率、鲁棒性与泛化能力方面均表现优异。

原文摘要 · Abstract (English)

Visual tokens consume substantial computational resources in multi-modal large models (MLLMs), significantly compromising their efficiency. Recent works have attempted to improve efficiency by compressing visual tokens during training, either through modifications to model components or by introducing additional parameters. However, they often overlook the increased learning difficulty caused by such compression, as the model's parameter space struggles to quickly adapt to the substantial perturbations in the feature space induced by token compression. In this work, we propose to develop Efficient MLLMs via Progressive Consistency Distillation (EPIC), a progressive learning framework. Specifically, by decomposing the feature space perturbations introduced by token compression along the token-wise and layer-wise dimensions, we introduce token consistency distillation and layer consistency distillation, respectively, aiming to reduce the training difficulty by leveraging guidance from a teacher model and following a progressive learning trajectory. Extensive experiments demonstrate the superior effectiveness, robustness, and generalization capabilities of our proposed framework.

多模态模型压缩蒸馏高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。