arXiv:2605.23482cs.CVcs.AI2026-05中稿 · CVPR被引 2

用几何感知方法高效压缩图文数据集,保持跨模态对齐。

Multimodal Distribution Matching for Vision-Language Dataset Distillation

论文配图:Multimodal Distribution Matching for Vision-Language Dataset Distillation
图 1 · 摘自论文原文
  • 在联合嵌入空间聚类中采样生成合成图文对
  • 通过加权插值混合教师模型,提升泛化性
  • 在单位超球面上匹配联合分布,适合多架构部署

数据集蒸馏可将大规模训练集压缩为紧凑的合成数据集,同时保留下游性能。随着多模态系统广泛使用配对的视觉-语言输入,多模态蒸馏需在有限算力和内存下保持表征质量与跨模态对齐,但现有方法常计算开销大且忽视模态间关联。为此,本文提出几何感知的多模态分布匹配(MDM)框架,实现高效通用的多模态蒸馏。MDM在数据、模型、损失三个层面集成互补组件:数据层面,从联合嵌入空间的聚类中采样初始化合成图像-文本对;模型层面,通过在权重空间按角度偏离预训练锚点的差异,插值独立微调的教师模型形成混合教师;损失层面,利用几何感知匹配目标,在单位超球面上匹配联合分布,结合跨模态一致与不一致方向的联合特征,并采用对称对比学习。在跨架构评估的图像-文本检索基准上,MDM生成的紧凑合成数据集能有效保持多模态语义,显著降低蒸馏成本,并在多种架构下保持鲁棒性。

原文摘要 · Abstract (English)

Dataset distillation compresses large training sets into compact synthetic datasets while preserving downstream performance. As modern systems increasingly operate on paired vision-language inputs, multimodal distillation must preserve representation quality and cross-modal alignment under tight compute and memory budgets, yet prior methods often require heavy computes and overlook their correlations. To address this, we present Multimodal Distribution Matching (MDM), a geometry-aware framework for efficient and generalizable multimodal distillation. Specifically, MDM integrates complementary components at the data, model, and loss levels. At the data level, it initializes synthetic image-text pairs by sampling from clusters in the joint embedding space. At the model level, it forms a mixed teacher by interpolating independently fine-tuned models in weight space according to their angular deviation from the pretrained anchor. At the loss level, it matches joint distributions on the unit hypersphere using a geometry-aware matching objective that exploits the joint features in the cross-modal agreement and discrepancy directions along with symmetric contrastive learning. Across image-text retrieval benchmarks with cross-architecture evaluation, MDM yields compact synthetic sets that preserve multimodal semantics, substantially reduce distillation cost, and remain robust across architectures.

数据蒸馏多模态视觉语言分布匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。