arXiv:2603.29239cs.CV2026-03

让扩散模型生成概念的清晰平均图像,揭示其内在语义结构。

Diffusion Mental Averages

  • 在扩散模型语义空间内对多条去噪轨迹进行对齐优化,生成平均原型。
  • 首次实现抽象概念的清晰、真实平均图像,避免传统方法的模糊结果。
  • 适用于多模态概念,可揭示模型偏见与概念表征,适合研究者使用。

扩散模型能否生成自身概念的‘心理平均’图像——即清晰、逼真的典型样本?我们提出扩散心理平均(DMA),一种模型中心的方法。以往基于数据集的平均方法在扩散模型生成样本上会产生模糊结果,且不考虑生成过程。相比之下,DMA 在扩散模型的语义空间中进行平均,该空间随时间演化且无直接解码器。我们将其建模为轨迹对齐:优化多个噪声隐变量,使其去噪轨迹逐步收敛至共享的粗粒度到细粒度语义,最终生成单一清晰原型。我们进一步将该方法扩展至多模态概念(如多种犬类),通过在 CLIP 等语义丰富空间聚类,并结合文本反转或 LoRA 将 CLIP 聚类映射至扩散空间。据我们所知,这是首个能稳定生成一致、真实平均图像的方法,即使对抽象概念也有效,可作为直观视觉摘要,亦可用于探究模型偏见与概念表征。

原文摘要 · Abstract (English)

Can a diffusion model produce its own "mental average" of a concept-one that is as sharp and realistic as a typical sample? We introduce Diffusion Mental Averages (DMA), a model-centric answer to this question. While prior methods aim to average image collections, they produce blurry results when applied to diffusion samples from the same prompt. These data-centric techniques operate outside the model, ignoring the generative process. In contrast, DMA averages within the diffusion model's semantic space, as discovered by recent studies. Since this space evolves across timesteps and lacks a direct decoder, we cast averaging as trajectory alignment: optimize multiple noise latents so their denoising trajectories progressively converge toward shared coarse-to-fine semantics, yielding a single sharp prototype. We extend our approach to multimodal concepts (e.g., dogs with many breeds) by clustering samples in semantically-rich spaces such as CLIP and applying Textual Inversion or LoRA to bridge CLIP clusters into diffusion space. This is, to our knowledge, the first approach that delivers consistent, realistic averages, even for abstract concepts, serving as a concrete visual summary and a lens into model biases and concept representation.

扩散模型概念平均语义空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。