arXiv:2510.17383cs.LGcs.CV2025-10

揭示生成模型隐空间演进,指出扩散模型打破统一表征假设。

The Evolving Nature of Latent Spaces: From GANs to Diffusion

  • 区分严格与广义合成:隐空间是否主导生成过程
  • 扩散模型将表征任务分散到各层,无统一隐空间
  • 适合研究生成机制、隐空间理论或媒体哲学的读者

本文考察生成视觉模型内部表征的演变,聚焦从GAN和VAE到扩散模型的范式转变。基于Beatrice Fazi关于合成是分布式表征整合的观点,提出区分'严格意义上的合成'(紧凑隐空间完全决定生成)与'广义合成'(表征劳动分布在各层)。通过剖析模型架构并设计干预层表示的实验,发现扩散模型将表征负担碎片化,挑战了统一内部空间的假设。结合媒体理论框架,批判性反思'隐空间'与'柏拉图表征假说'等隐喻,主张重新理解生成AI:并非直接合成内容,而是由专门化过程涌现而成的配置。

原文摘要 · Abstract (English)

This paper examines the evolving nature of internal representations in generative visual models, focusing on the conceptual and technical shift from GANs and VAEs to diffusion-based architectures. Drawing on Beatrice Fazi's account of synthesis as the amalgamation of distributed representations, we propose a distinction between "synthesis in a strict sense", where a compact latent space wholly determines the generative process, and "synthesis in a broad sense," which characterizes models whose representational labor is distributed across layers. Through close readings of model architectures and a targeted experimental setup that intervenes in layerwise representations, we show how diffusion models fragment the burden of representation and thereby challenge assumptions of unified internal space. By situating these findings within media theoretical frameworks and critically engaging with metaphors such as the latent space and the Platonic Representation Hypothesis, we argue for a reorientation of how generative AI is understood: not as a direct synthesis of content, but as an emergent configuration of specialized processes.

隐空间生成模型扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。