arXiv:2603.25388cs.CV2026-03中稿 · ICLR被引 1

用分阶段教师模型提升多模态数据蒸馏效果

Multimodal Dataset Distillation via Phased Teacher Models

  • 分阶段建模教师学习动态,用捷径路径稳定蒸馏过程
  • 在Flickr30k上最高提升13.5%,平均增益9.53%
  • 适合关注高效多模态知识压缩的研究者

多模态数据蒸馏旨在构建紧凑的合成数据集,实现大规模图像-文本数据的高效压缩与知识迁移。然而,现有方法难以捕捉教师模型后期训练阶段中复杂的动态知识,导致学生模型性能下降和蒸馏数据质量受损。为应对显著的跨阶段性能差距与不稳定的教师轨迹问题,我们提出分阶段教师模型与捷径轨迹(PTM-ST)——一种新型分阶段蒸馏框架。PTM-ST通过阶段感知的教师建模与基于捷径的轨迹构建策略,精确拟合教师在不同训练阶段的学习动态,提升蒸馏过程的稳定性和表达能力。理论分析与全面实验表明,PTM-ST显著缓解优化震荡与阶段间知识鸿沟,同时降低存储开销。该方法在Flickr30k与COCO数据集上持续优于现有最先进基线,其中在Flickr30k上达到最高13.5%的绝对提升,平均增益9.53%。

原文摘要 · Abstract (English)

Multimodal dataset distillation aims to construct compact synthetic datasets that enable efficient compression and knowledge transfer from large-scale image-text data. However, existing approaches often fail to capture the complex, dynamically evolving knowledge embedded in the later training stages of teacher models. This limitation leads to degraded student performance and compromises the quality of the distilled data. To address critical challenges such as pronounced cross-stage performance gaps and unstable teacher trajectories, we propose Phased Teacher Model with Shortcut Trajectory (PTM-ST) -- a novel phased distillation framework. PTM-ST leverages stage-aware teacher modeling and a shortcut-based trajectory construction strategy to accurately fit the teacher's learning dynamics across distinct training phases. This enhances both the stability and expressiveness of the distillation process. Through theoretical analysis and comprehensive experiments, we show that PTM-ST significantly mitigates optimization oscillations and inter-phase knowledge gaps, while also reducing storage overhead. Our method consistently surpasses state-of-the-art baselines on Flickr30k and COCO, achieving up to 13.5% absolute improvement and an average gain of 9.53% on Flickr30k. Code: https://github.com/Previsior/PTM-ST.

数据蒸馏多模态教师模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。