arXiv:2606.03345cs.CVcs.CL2026-06

用视觉语言数据建模跨文化感知体验,区分事实与情感维度。

Beyond Semantics: Modeling Factual and Affective Perceptual Experiences from Vision-Language Data

论文配图:Beyond Semantics: Modeling Factual and Affective Perceptual Experiences from Vision-Language Data
图 1 · 摘自论文原文
  • 通过无监督聚类发现图像的感知主题,自动确定最优聚类数。
  • 在ArtELingo数据集上轮廓系数达0.97,映射准确率AUC为0.94。
  • 能捕捉有意义的感知差异,适合跨文化视觉理解研究者使用。

我们提出P-Topics(感知主题)建模,旨在理解图像在不同文化和情感层面的感知方式。目标是(1)从图像与文字配对数据中发现并建模多种感知体验,每种体验包含客观事实和主观情感两个方面;(2)将图像与对应的感知体验关联。为此,我们设计了PercepT(Perception Topic Transformer),一种两阶段架构:第一阶段通过无监督训练目标发现P-Topics作为视觉-文本聚类,并动态选择聚类数量以匹配数据的感知丰富性;第二阶段利用注意力池化学习映射函数,将图像关联到相应聚类。在ArtELingo数据集上,PercepT的轮廓系数达到0.97,远超基线的0.37;映射任务的AUC达0.94,优于基线的0.77。人工评估确认其能捕捉语义上合理的感知体验,显著优于现有方法。代码将公开。

原文摘要 · Abstract (English)

We present P-Topics (Perception Topics) modeling, a novel problem for understanding how images are perceived affectively and across cultures. The goal is to (1) discover and model the different perception experiences in a dataset of images and captions, where each experience is defined by an objective factual and a subjective affective aspect, and (2) associate images to their relevant perception experiences. We introduce **PercepT** (**Percep**tion topic **T**ransformer), a two-stage architecture that tackles P-Topics modeling. In the formation stage, percepT discovers *P-Topics* as visual-textual clusters using an unsupervised training objective, and dynamically selects the number of clusters to match the perceptual richness of the dataset. In the mapping stage, it learns *P-Topic mapping functions* via attention pooling to associate images to their respective clusters. On ArtELingo, PercepT achieves a silhouette score of **0.97** compared to **0.37** from the closest baseline reflecting better perceptual clusters. PercepT also achieves an AUC score of **0.94** compared to **0.77** showing better mapping to perceptual clusters. Human evaluation confirms that PercepT captures semantically meaningful perception experiences and significantly outperforms existing methods. Our implementation will be made public.

感知建模多模态跨文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。