arXiv:2608.30366cs.LG2026-08

首次发现生成与对比模型的模式连通性,揭示其损失景观的几何特性。

Mode Connectivity Beyond Classifiers: Evidence from Generative and Contrastive Models

论文配图:Mode Connectivity Beyond Classifiers: Evidence from Generative and Contrastive Models
图 1 · 摘自论文原文
  • 针对扩散模型和对比学习模型设计专用连接算法
  • 首次在独立训练的DDPM与NanoCLIP间发现低损耗连续路径
  • 为理解生成与对比模型的优化结构提供新视角

深度神经网络的损失景观具有高度复杂且非凸的特性。近期研究揭示了模式连通性现象,即独立训练的网络模式可通过连续低损耗路径相连。然而,现有研究主要局限于分类器模型,现代复杂模型中是否存在类似几何性质仍不清楚。本文将模式连通性拓展至生成与对比学习领域(具体为DDPM和NanoCLIP),针对其独特架构提出一种架构感知的连接构建算法。大量实证结果首次证明,可成功发现独立训练的DDPM与NanoCLIP模式间的模式连通性。本工作为理解现代生成与对比模型损失景观的几何特性提供了新视角。

原文摘要 · Abstract (English)

The loss landscape of Deep Neural Networks (DNNs) exhibits highly complex and non-convex properties. Recent studies have revealed the phenomenon of mode connectivity, demonstrating that independently trained network modes can be connected via a continuous low-loss path. However, existing mode connectivity research is predominantly confined to classifier-based models, leaving it an open question whether similar geometric properties exist in modern complex models. In this paper, we extend the boundaries of mode connectivity to generative and contrastive domains (specifically DDPM and NanoCLIP). Addressing the unique architecture of DDPM and CLIP, we propose an architecture-aware connection building algorithm. Extensive empirical results demonstrate for the first time that we successfully discover mode connectivity between independently trained DDPM and NanoCLIP modes. Our work provides a novel perspective for understanding the geometric properties of the loss landscapes in modern generative and contrastive models.

模式连通性生成模型对比学习损失景观

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。