提出首个多模态对比学习记忆度量方法,可有效抑制模型过拟合噪声。
MultiMem: Measuring and Mitigating Memorization in Multi-Modal Contrastive Learning

- 设计新指标MultiMem,量化多模态对比学习中的记忆程度。
- 发现文本模态最易引发记忆,跨模态语义错位影响最大。
- 通过跨模态增强可显著降低记忆,提升模型泛化性能。
机器学习模型的记忆能力虽能提升对罕见样本的性能,但也导致噪声和异常值的有害保留,损害泛化性。尽管在视觉领域的监督与自监督学习中已广泛研究记忆现象,但多模态对比学习中的记忆问题仍待探索。本文提出MultiMem,首个用于量化多模态对比学习记忆程度的指标。系统分析表明,跨模态语义错位对记忆影响最强,其中文本模态主导记忆行为,其次为视频、图像和音频。实验显示,对所有模态施加针对性增强可有效降低由MultiMem衡量的记忆水平,并提升模型性能。本工作建立了首个多模态对比学习中测量与缓解记忆的框架,防止有害数据留存,助力构建更优模型。
原文摘要 · Abstract (English)
Memorization in machine learning models enables high performance on rare in-distribution samples by capturing their atypical patterns. However, it also causes harmful retention of noise and outliers, degrading generalization. While memorization has been extensively studied in both supervised and self-supervised learning in the vision domain, it remains unexplored in multi-modal contrastive learning. We address this gap by introducing MultiMem, the first metric designed to quantify memorization in multi-modal contrastive learning. Through our systematic analysis, we demonstrate that cross-modal semantic misalignment has the strongest influence on memorization, with text being the dominant modality driving memorization, followed by video, image, and audio. We show that targeted augmentations applied across all modalities effectively reduce memorization as measured by our MultiMem metric and improve model performance. Overall, this work establishes the first framework for measuring and mitigating memorization in multi-modal contrastive learning, preventing harmful data retention and contributing to higher-performing models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。