arXiv:2502.19642cs.LG2025-02

提出cMIM框架,统一生成与判别任务的表征学习。

Contrastive MIM: A Contrastive Mutual Information Framework for Unified Generative and Discriminative Representation Learning

  • 引入对比目标增强MIM,无需正样本增强且对批大小鲁棒
  • 在分类与回归任务中优于MIM和InfoNCE,重建质量相当
  • 提出信息嵌入技术,无需训练即可提升判别性能

学习能泛化到未知下游任务的表征是表示学习的核心挑战。现有方法如对比学习、自监督掩码和去噪自编码器各有权衡。本文提出对比互信息机(cMIM),一种基于概率框架的MIM扩展,通过新颖的对比目标提升表征能力。虽然MIM能最大化输入与隐变量间的互信息并促进隐码聚类,但在判别任务上表现逊于先进方法。cMIM在保持MIM生成优势的同时,强制全局判别结构,解决此局限。主要贡献有二:(1) 提出cMIM,无需正样本增强,对批大小鲁棒,不同于InfoNCE;(2) 引入信息嵌入,从编码器-解码器模型中提取丰富表征,显著提升判别性能,且无需额外训练,适用范围广。实验证明,cMIM在分类与回归任务中持续优于MIM和InfoNCE,同时保持相当的重建质量。结果表明,cMIM为生成与判别应用提供了统一有效的表征学习框架。

原文摘要 · Abstract (English)

Learning representations that generalize well to unknown downstream tasks is a central challenge in representation learning. Existing approaches such as contrastive learning, self-supervised masking, and denoising auto-encoders address this challenge with varying trade-offs. In this paper, we introduce the {contrastive Mutual Information Machine} (cMIM), a probabilistic framework that augments the Mutual Information Machine (MIM) with a novel contrastive objective. While MIM maximizes mutual information between inputs and latent variables and encourages clustering of latent codes, its representations underperform on discriminative tasks compared to state-of-the-art alternatives. cMIM addresses this limitation by enforcing global discriminative structure while retaining MIM's generative strengths. We present two main contributions: (1) we propose cMIM, a contrastive extension of MIM that eliminates the need for positive data augmentation and is robust to batch size, unlike InfoNCE-based methods; (2) we introduce {informative embeddings}, a general technique for extracting enriched representations from encoder--decoder models that substantially improve discriminative performance without additional training, and which apply broadly beyond MIM. Empirical results demonstrate that cMIM consistently outperforms MIM and InfoNCE in classification and regression tasks, while preserving comparable reconstruction quality. These findings suggest that cMIM provides a unified framework for learning representations that are simultaneously effective for discriminative and generative applications.

表征学习对比学习生成模型互信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。