arXiv:2505.14705cs.CVcs.LG2025-05NeurIPS被引 9

解决多模态数据蒸馏中的模态坍缩问题,提升跨模态学习效果。

Beyond Modality Collapse: Representations Blending for Multimodal Dataset Distillation

  • 通过表示融合削弱主导的跨模态监督,增强模态内多样性。
  • 在Flickr-30K和MS-COCO上实现+9.4(IR@10)检索性能提升。
  • 适合关注多模态模型压缩与高效训练的研究者。

多模态数据蒸馏(MDD)旨在将大规模图像-文本数据集压缩为紧凑代理数据集,同时保持其跨模态学习的有效性。尽管近期取得进展,现有MDD方法常面临模态坍缩问题,表现为模态内表示过度集中、跨模态分布差距扩大。本文首次指出该问题源于数据蒸馏固有的过压缩行为与对比学习目标之间的根本冲突。为此,提出RepBlend框架,通过表示融合弱化过强的跨模态监督,显著提升模态内多样性。此外,发现现有方法存在模态间不对称监督导致优化偏差,提出对称投影轨迹匹配策略,利用模态特定投影头同步优化动态,促进平衡监督与跨模态对齐。在Flickr-30K和MS-COCO上的实验表明,RepBlend持续优于现有最先进MDD方法,在100对设置下检索性能提升达+9.4(IR@10)和+6.3(TR@10),并实现最高6.7×的数据蒸馏加速。

原文摘要 · Abstract (English)

Multimodal Dataset Distillation (MDD) seeks to condense large-scale image-text datasets into compact surrogates while retaining their effectiveness for cross-modal learning. Despite recent progress, existing MDD approaches often suffer from \textit{\textbf{Modality Collapse}}, characterized by over-concentrated intra-modal representations and enlarged distributional gap across modalities. In this paper, at the first time, we identify this issue as stemming from a fundamental conflict between the over-compression behavior inherent in dataset distillation and the cross-modal supervision imposed by contrastive objectives. To alleviate modality collapse, we introduce \textbf{RepBlend}, a novel MDD framework that weakens overdominant cross-modal supervision via representation blending, thereby significantly enhancing intra-modal diversity. Additionally, we observe that current MDD methods impose asymmetric supervision across modalities, resulting in biased optimization. To address this, we propose symmetric projection trajectory matching, which synchronizes the optimization dynamics using modality-specific projection heads, thereby promoting balanced supervision and enhancing cross-modal alignment. Experiments on Flickr-30K and MS-COCO show that RepBlend consistently outperforms prior state-of-the-art MDD methods, achieving significant gains in retrieval performance (e.g., +9.4 IR@10, +6.3 TR@10 under the 100-pair setting) and offering up to 6.7$\times$ distillation speedup.

多模态数据蒸馏表示学习对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。