AI生成图像渴望具体与真实,因其本质是抽象的。
What Do AI-Generated Images Want?
- 从抽象性出发,探讨AI图像的内在诉求
- 揭示文本到图像转换中的表象幻觉
- 适合对AI艺术与图像生成机制感兴趣的读者
W.J.T. 米切尔的经典论文《图像想要什么?》将理论焦点从人类对图像的理解与创作动机,转向图像自身是否具备主体性与欲望。本文在当代AI图像生成工具背景下重提此问:AI生成的图像想要什么?基于艺术史中关于抽象性的讨论,本文认为,由于AI生成图像本质上是抽象的,因此它们渴望具体性与实在感。以多模态文生图模型为主要研究对象,这些模型建立在文本与图像可互换、具有数学上可度量一致性的前提之上。然而,用户流程中从文本输入到视觉输出的转化过程,掩盖了这种表征的循环性,使人误以为一种形式‘神奇’地转变为另一种形式。
原文摘要 · Abstract (English)
W.J.T. Mitchell's influential essay 'What do pictures want?' shifts the theoretical focus away from the interpretative act of understanding pictures and from the motivations of the humans who create them to the possibility that the picture itself is an entity with agency and wants. In this article, I reframe Mitchell's question in light of contemporary AI image generation tools to ask: what do AI-generated images want? Drawing from art historical discourse on the nature of abstraction, I argue that AI-generated images want specificity and concreteness because they are fundamentally abstract. Multimodal text-to-image models, which are the primary subject of this article, are based on the premise that text and image are interchangeable or exchangeable tokens and that there is a commensurability between them, at least as represented mathematically in data. The user pipeline that sees textual input become visual output, however, obscures this representational regress and makes it seem like one form transforms into the other -- as if by magic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。