解析生成模型潜空间的双重属性:文化档案与潜在可能性。
Reflections on Disentanglement and the Latent Space
- 将潜空间视为文化记忆库与潜在能力场的双重视角。
- 通过解耦机制揭示潜空间的组织逻辑,使其更易人类理解。
- 关联德勒兹潜能哲学与休谟想象力理论,拓展艺术与技术对话。
图像生成模型的潜空间是压缩的隐含视觉知识多维空间,吸引计算机科学家、数字艺术家与媒体学者关注。该空间已成为人工智能艺术中的审美范畴,催生如马里奥·克林格曼等人的潜空间漫步等创作手法。它也被视为视觉文化的缩影,编码了丰富的视觉世界表征。本文提出潜空间的双重观:既是多维文化档案,也是多维潜能空间。论文探讨解耦作为阐明这一双重性的方法,并将其作为以人类可理解方式利用空间结构的解释路径。文章对比解耦作为潜能与条件化作为想象的作用,结合德勒兹潜能哲学与休谟想象力理论进行阐释。最后指出传统生成模型与近期架构之间的差异。
原文摘要 · Abstract (English)
The latent space of image generative models is a multi-dimensional space of compressed hidden visual knowledge. Its entity captivates computer scientists, digital artists, and media scholars alike. Latent space has become an aesthetic category in AI art, inspiring artistic techniques such as the latent space walk, exemplified by the works of Mario Klingemann and others. It is also viewed as cultural snapshots, encoding rich representations of our visual world. This paper proposes a double view of the latent space, as a multi-dimensional archive of culture and as a multi-dimensional space of potentiality. The paper discusses disentanglement as a method to elucidate the double nature of the space and as an interpretative direction to exploit its organization in human terms. The paper compares the role of disentanglement as potentiality to that of conditioning, as imagination, and confronts this interpretation with the philosophy of Deleuzian potentiality and Hume's imagination. Lastly, this paper notes the difference between traditional generative models and recent architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。