arXiv:2506.01586cs.CVcs.LG2025-06被引 1

从嘈杂的多模态数据中提炼出高效干净的数据集,提升模型训练效率。

Multi-Modal Dataset Distillation in the Wild

  • 通过可学习细粒度对应关系,自适应优化关键区域,提升数据信息密度。
  • 在多种压缩比下性能优于现有方法超15%,显著降低资源消耗。
  • 适合需要高效训练且数据质量参差的多模态应用开发者使用。

近年来多模态模型在真实场景中展现出强大泛化能力,但其快速发展面临两大数据挑战:一是训练需大规模数据,带来高昂存储与计算成本;二是数据多为网络爬取,存在不可避免的噪声(如部分不匹配对),严重损害模型性能。为此,我们提出首个面向野外复杂场景的多模态数据蒸馏框架——MDW。MDW通过引入可学习的细粒度跨模态对应关系,并自适应优化蒸馏数据以强调具有判别性的对应区域,从而提升数据的信息密度与训练有效性。此外,为从真实数据中获取稳健的跨模态先验知识,提出双轨协同学习机制,避免噪声干扰,实现可验证的噪声容忍度。大量实验验证了MDW在理论与实践上的有效性,具备优异扩展性,在多种压缩比下均超越现有方法超15%,展现出在多样效能与资源需求场景下的实用潜力。

原文摘要 · Abstract (English)

Recent multi-modal models have shown remarkable versatility in real-world applications. However, their rapid development encounters two critical data challenges. First, the training process requires large-scale datasets, leading to substantial storage and computational costs. Second, these data are typically web-crawled with inevitable noise, i.e., partially mismatched pairs, severely degrading model performance. To these ends, we propose Multi-modal dataset Distillation in the Wild, i.e., MDW, the first framework to distill noisy multi-modal datasets into compact clean ones for effective and efficient model training. Specifically, MDW introduces learnable fine-grained correspondences during distillation and adaptively optimizes distilled data to emphasize correspondence-discriminative regions, thereby enhancing distilled data's information density and efficacy. Moreover, to capture robust cross-modal correspondence prior knowledge from real data, MDW proposes dual-track collaborative learning to avoid the risky data noise, alleviating information loss with certifiable noise tolerance. Extensive experiments validate MDW's theoretical and empirical efficacy with remarkable scalability, surpassing prior methods by over 15% across various compression ratios, highlighting its appealing practicality for applications with diverse efficacy and resource needs.

多模态数据蒸馏噪声容忍高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。