arXiv:2604.11374cs.CVcs.CL2026-04ACL被引 1

探索视觉语言模型如何编码个人审美偏好,实现无需微调的轻量级个性化图像评分。

What Do Vision-Language Models Encode for Personalized Image Aesthetics Assessment?

  • 分析模型内部表示,发现美学特征分布在语言解码层
  • 仅用线性模型即可实现有效个性化图像评分
  • 适用于无须微调的快速个性审美建模

个性化图像美学评估(PIAA)在实际应用中具有重要意义。尽管基于视觉语言模型(VLMs)的方法在该任务上表现出潜力,但其内部是否编码了丰富且多层次的美学属性仍不明确。本文首先分析VLM的内部表征,考察美学属性的存在与分布,并在此基础上实现无需模型微调的轻量级个体化评估。分析表明,VLM在语言解码层中编码了多样化的美学特征。基于这些表征,我们证明简单的线性模型即可有效完成PIAA。此外,我们还研究了不同VLM架构及图像领域间美学信息的跨层传递规律。结果为利用VLM建模主观、个体化的审美偏好提供了新见解。代码已公开于 https://github.com/ynklab/vlm-latent-piaa。

原文摘要 · Abstract (English)

Personalized image aesthetics assessment (PIAA) is an important research problem with practical real-world applications. While methods based on vision-language models (VLMs) are promising candidates for PIAA, it remains unclear whether they internally encode rich, multi-level aesthetic attributes required for effective personalization. In this paper, we first analyze the internal representations of VLMs to examine the presence and distribution of such aesthetic attributes, and then leverage them for lightweight, individual-level personalization without model fine-tuning. Our analysis reveals that VLMs encode diverse aesthetic attributes that propagate into the language decoder layers. Building on these representations, we demonstrate that simple linear models can perform PIAA effectively. We further analyze how aesthetic information is transferred across layers in different VLM architectures and across image domains. Our findings provide insights into how VLMs can be utilized for modeling subjective, individual aesthetic preferences. Our code is available at https://github.com/ynklab/vlm-latent-piaa.

视觉语言模型个性化评估美学分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。