arXiv:2509.21511cs.LGstat.ML2025-09被引 1

无需正样本增强,用对比学习提升模型判别能力

Contrastive Mutual Information Learning: Toward Robust Representations without Positive-Pair Augmentations

  • 提出cMIM框架,用对比目标替代正样本增强
  • 在图像和分子数据上分类回归性能优于MIM和InfoNCE
  • 可直接提取丰富特征,不需额外训练,适用广泛

学习能泛化到多样下游任务的表示仍是表示学习的核心挑战。现有范式——对比学习、自监督掩码和去噪自编码器——在这一挑战上各有权衡。我们提出对比互信息机(cMIM),一个基于概率框架的扩展,将互信息机(MIM)与对比目标结合。虽然MIM最大化输入与潜在表示间的互信息并促进代码聚类,但在判别任务上表现不足。cMIM通过引入全局判别结构,在保持MIM生成保真度的同时弥补了这一缺陷。贡献有三:首先,提出cMIM,一种无需正样本增强且对批次大小不敏感的MIM对比扩展;其次,提出‘信息嵌入’技术,从编码器-解码器模型中提取增强特征,无需额外训练即可提升判别性能,并可广泛应用于MIM之外;第三,通过视觉与分子基准的实证表明,cMIM在分类与回归任务上优于MIM和InfoNCE,同时保持竞争力的重建质量。这些结果使cMIM成为统一的表示学习框架,推动模型在判别与生成应用中高效协同。

原文摘要 · Abstract (English)

Learning representations that transfer well to diverse downstream tasks remains a central challenge in representation learning. Existing paradigms -- contrastive learning, self-supervised masking, and denoising auto-encoders -- balance this challenge with different trade-offs. We introduce the {contrastive Mutual Information Machine} (cMIM), a probabilistic framework that extends the Mutual Information Machine (MIM) with a contrastive objective. While MIM maximizes mutual information between inputs and latents and promotes clustering of codes, it falls short on discriminative tasks. cMIM addresses this gap by imposing global discriminative structure while retaining MIM's generative fidelity. Our contributions are threefold. First, we propose cMIM, a contrastive extension of MIM that removes the need for positive data augmentation and is substantially less sensitive to batch size than InfoNCE. Second, we introduce {informative embeddings}, a general technique for extracting enriched features from encoder-decoder models that boosts discriminative performance without additional training and applies broadly beyond MIM. Third, we provide empirical evidence across vision and molecular benchmarks showing that cMIM outperforms MIM and InfoNCE on classification and regression tasks while preserving competitive reconstruction quality. These results position cMIM as a unified framework for representation learning, advancing the goal of models that serve both discriminative and generative applications effectively.

对比学习表示学习无监督学习互信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。