揭示视觉语言模型的潜在空间并非中立,而是带有特定政治与文化偏见。
Ethology of Latent Spaces
- 通过对比三类模型在301幅艺术作品上的表现,发现其对政治性与美学的判断差异显著。
- SigLIP认为59.4%的艺术品具政治性,而OpenCLIP仅4%;非洲面具在不同模型中评分差距达72.6个百分点。
- 提出算法可见性三范式:熵、制度与符号,适用于批判性分析AI对艺术文化的解读。
本研究从行为学视角挑战视觉语言模型(VLMs)潜在空间固有中立性的假设。不同于均匀无差别的空间,潜在空间展现出由训练数据与架构决定的模型特异性算法敏感性,表现为感知显著性的不同调节机制。通过对301件(15至20世纪)艺术作品应用三种模型(OpenAI CLIP、OpenCLIP LAION、SigLIP)进行对比分析,揭示其在政治与文化类别归属上存在显著分歧。基于向量类比(Mikolov et al., 2013)构建双极语义轴,发现SigLIP将59.4%的作品判定为具有政治性,而OpenCLIP仅为4%;非洲面具在SigLIP中得分最高,在OpenAI CLIP中则被视为无政治性。在美学殖民轴上,模型间差异高达72.6个百分点。本文提出三个操作概念:计算潜在政治化(computational latent politicization)、涌现偏差(emergent bias),以及三种算法视野范式——熵型(LAION)、制度型(OpenAI)、符号型(SigLIP),分别塑造不同的可见性模式。结合福柯的档案概念、詹姆斯翁的意识形态原型及西蒙东的个体化理论,认为训练数据集如同准档案,其话语形式在潜在空间中结晶。该研究呼吁对VLM在数字艺术史中的应用进行批判性重估,并倡导将学习架构纳入文化解释权委托的考量之中。
原文摘要 · Abstract (English)
This study challenges the presumed neutrality of latent spaces in vision language models (VLMs) by adopting an ethological perspective on their algorithmic behaviors. Rather than constituting spaces of homogeneous indeterminacy, latent spaces exhibit model-specific algorithmic sensitivities, understood as differential regimes of perceptual salience shaped by training data and architectural choices. Through a comparative analysis of three models (OpenAI CLIP, OpenCLIP LAION, SigLIP) applied to a corpus of 301 artworks (15th to 20th), we reveal substantial divergences in the attribution of political and cultural categories. Using bipolar semantic axes derived from vector analogies (Mikolov et al., 2013), we show that SigLIP classifies 59.4% of the artworks as politically engaged, compared to only 4% for OpenCLIP. African masks receive the highest political scores in SigLIP while remaining apolitical in OpenAI CLIP. On an aesthetic colonial axis, inter-model discrepancies reach 72.6 percentage points. We introduce three operational concepts: computational latent politicization, describing the emergence of political categories without intentional encoding; emergent bias, irreducible to statistical or normative bias and detectable only through contrastive analysis; and three algorithmic scopic regimes: entropic (LAION), institutional (OpenAI), and semiotic (SigLIP), which structure distinct modes of visibility. Drawing on Foucault's notion of the archive, Jameson's ideologeme, and Simondon's theory of individuation, we argue that training datasets function as quasi-archives whose discursive formations crystallize within latent space. This work contributes to a critical reassessment of the conditions under which VLMs are applied to digital art history and calls for methodologies that integrate learning architectures into any delegation of cultural interpretation to algorithmic agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。