arXiv:2510.14190cs.LG2025-10中稿 · ICML被引 1

用对比学习让扩散模型的隐空间更可控,支持平滑编辑与推理。

Contrastive Diffusion Alignment: Learning Structured Latents for Controllable Generation

  • 通过对比学习构建低维隐空间,方向对齐动态因素。
  • 在多种数据上实现更可解释的平滑插值与反事实编辑。
  • 适用于流体、神经信号等多领域,兼容不同反演方法。

扩散模型生成能力强,但其隐空间维度高且缺乏显式结构,难以解释和控制。我们提出ConDA(对比扩散对齐),一种即插即用的几何层,利用辅助变量(如时间、刺激参数、面部动作单元)对预训练扩散模型的隐变量施加对比学习。ConDA学习到一个低维嵌入,其方向与潜在动态因素对齐,符合近期关于结构化、解耦表示的对比学习成果。在此嵌入中,简单的非线性轨迹支持平滑插值、外推和反事实编辑,而渲染仍保持在原始扩散空间。ConDA通过保留邻域关系的kNN解码器将嵌入轨迹映射回扩散隐变量,实现编辑与渲染分离,并对不同反演求解器具有鲁棒性。在流体动力学、神经钙成像、治疗性神经刺激、面部表情动态及猴运动皮层活动等多个场景中,ConDA生成的隐结构比线性遍历和基于条件的基线更具可解释性和可控性,表明扩散隐变量编码了可被显式对比几何层挖掘的动力学相关结构。

原文摘要 · Abstract (English)

Diffusion models excel at generation, but their latent spaces are high dimensional and not explicitly organized for interpretation or control. We introduce ConDA (Contrastive Diffusion Alignment), a plug-and-play geometry layer that applies contrastive learning to pretrained diffusion latents using auxiliary variables (e.g., time, stimulation parameters, facial action units). ConDA learns a low-dimensional embedding whose directions align with underlying dynamical factors, consistent with recent contrastive learning results on structured and disentangled representations. In this embedding, simple nonlinear trajectories support smooth interpolation, extrapolation, and counterfactual editing while rendering remains in the original diffusion space. ConDA separates editing and rendering by lifting embedding trajectories back to diffusion latents with a neighborhood-preserving kNN decoder and is robust across inversion solvers. Across fluid dynamics, neural calcium imaging, therapeutic neurostimulation, facial expression dynamics, and monkey motor cortex activity, ConDA yields more interpretable and controllable latent structure than linear traversals and conditioning-based baselines, indicating that diffusion latents encode dynamics-relevant structure that can be exploited by an explicit contrastive geometry layer.

扩散模型可控生成对比学习隐空间结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。