测试CLIP理解艺术与人类是否一致,发现它有优势也有盲区。
Does CLIP perceive art the same way we do?
- 设计多种探测任务评估CLIP对画作内容、风格等的理解能力。
- 在艺术风格和历史时期识别上表现良好,但对审美意图感知不足。
- 适合研究生成模型中的视觉语义对齐,尤其关注创意领域应用。
CLIP作为一种强大的多模态模型,能通过联合嵌入连接图像与文本,但其'看'艺术的方式与人类有多接近?本文探究了CLIP从绘画(包括人类创作与AI生成)中提取高层次语义与风格信息的能力,涵盖内容、场景理解、艺术风格、历史时期及视觉畸变或伪影等维度。通过设计针对性探测任务,并对比人类标注与专家基准,分析其与人类感知和上下文理解的一致性。结果揭示了CLIP视觉表征在美学线索与艺术意图方面的优势与局限。研究进一步讨论了这些发现对生成过程中使用CLIP作为引导机制(如风格迁移或基于提示的图像合成)的影响。工作强调了在创造性领域应用时,多模态系统需具备更深的可解释性。
原文摘要 · Abstract (English)
CLIP has emerged as a powerful multimodal model capable of connecting images and text through joint embeddings, but to what extent does it 'see' the same way humans do - especially when interpreting artworks? In this paper, we investigate CLIP's ability to extract high-level semantic and stylistic information from paintings, including both human-created and AI-generated imagery. We evaluate its perception across multiple dimensions: content, scene understanding, artistic style, historical period, and the presence of visual deformations or artifacts. By designing targeted probing tasks and comparing CLIP's responses to human annotations and expert benchmarks, we explore its alignment with human perceptual and contextual understanding. Our findings reveal both strengths and limitations in CLIP's visual representations, particularly in relation to aesthetic cues and artistic intent. We further discuss the implications of these insights for using CLIP as a guidance mechanism during generative processes, such as style transfer or prompt-based image synthesis. Our work highlights the need for deeper interpretability in multimodal systems, especially when applied to creative domains where nuance and subjectivity play a central role.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。