arXiv:2608.27121cs.MMcs.CV2026-08中稿 · ICMI Companion '26

AI可自发发现艺术美学结构,无需人工标注。

How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space

论文配图:How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space
图 1 · 摘自论文原文
  • 自监督框架将四模态数据映射到256维共享空间
  • 通过迭代聚类发现隐含美学分组,与人类情感标签存在差异
  • 适合研究AI理解艺术、媒体组织与自动标注的学者

美学是艺术作品象征意义的重要部分。尽管主观,人类仍能基于引发的情绪对不同模态的艺术作品进行分类。目前尚不清楚的是,人工智能模型在无显式标签或跨模态监督的情况下如何形成对人类创作媒介的审美分类。本文提出一种自监督框架,将文本、音频、图像和视频四种模态投影至一个256维共享嵌入空间,并采用迭代聚类方法挖掘美学结构。我们在弱监督多模态数据集上分析了AI生成聚类结果与人类情感标注之间的差异。该研究在理解AI如何构建跨模态相似性、组织异构媒体集合用于检索增强生成(RAG),以及自动化数据标注方面具有应用价值。

原文摘要 · Abstract (English)

Aesthetics are an important part of the symbolism of artistic works. Although subjective, humans categorize art based on the emotion evoked regardless of modality. What remains under-explored is how AI models form their own aesthetic categorization of human-produced media without explicit labels or cross-modal supervision. We present a self-supervised framework that projects four modalities (text, audio, image and video) into a shared 256-dimensional embedding space and applies iterative clustering to discover aesthetic structure. We discuss the divergence between AI-generated cluster assignments and human affective register labels on a weakly supervised multimodal dataset. This work has applications in understanding how AI structures cross-modal similarity, organizing heterogeneous media collections for Retrieval-Augmented Generation (RAG), and automated data labeling.

美学自监督多模态嵌入空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。