arXiv:2411.18674cs.CVcs.LG2024-11CVPR被引 22

通过主动筛选数据提升多模态模型压缩效果,比传统方法更高效。

Active Data Curation Effectively Distills Large-Scale Multimodal Models

  • 用在线批次选择法主动挑好数据,代替复杂蒸馏策略。
  • 在27个零样本任务上达顶尖性能,推理计算量减少最多11%。
  • 适合需要高效推理的多模态应用,如图文生成与问答系统。

知识蒸馏(KD)是将大规模模型压缩为小型模型的标准方法。以往研究探索了越来越复杂的蒸馏策略,包括不同目标函数、教师模型集成和权重继承。本文提出一种替代性但简洁的方法——通过主动数据筛选实现对比式多模态预训练的有效蒸馏。我们提出的在线批次选择方法ACID,在多种模型、数据与计算配置下均优于强基线。进一步发现,该策略与标准KD具有互补性,可有效结合以训练高性能且推理高效的模型。我们的简单可扩展预训练框架ACED,在27个零样本分类与检索任务中达到当前最优表现,推理浮点运算量最多降低11%。此外,我们在LiT-Decoder设置下证明,ACED模型生成的视觉编码器在训练生成式多模态模型时表现优异,超越更大视觉编码器在图像描述与视觉问答任务中的表现。

原文摘要 · Abstract (English)

Knowledge distillation (KD) is the de facto standard for compressing large-scale models into smaller ones. Prior works have explored ever more complex KD strategies involving different objective functions, teacher-ensembles, and weight inheritance. In this work we explore an alternative, yet simple approach -- active data curation as effective distillation for contrastive multimodal pretraining. Our simple online batch selection method, ACID, outperforms strong KD baselines across various model-, data- and compute-configurations. Further, we find such an active data curation strategy to in fact be complementary to standard KD, and can be effectively combined to train highly performant inference-efficient models. Our simple and scalable pretraining framework, ACED, achieves state-of-the-art results across 27 zero-shot classification and retrieval tasks with upto 11% less inference FLOPs. We further demonstrate that our ACED models yield strong vision-encoders for training generative multimodal models in the LiT-Decoder setting, outperforming larger vision encoders for image-captioning and visual question-answering tasks.

知识蒸馏多模态高效推理数据筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。