研究抽象与具体概念的视觉表现差异,发现简单特征更有效区分类别。
Unveiling the Mystery of Visual Attributes of Concrete and Abstract Concepts: Variability, Nearest Neighbors, and Challenging Categories
- 用1000个概念图像分析视觉多样性,对比抽象与具体概念
- 颜色纹理等基础特征比ViT在分类中表现更好,但ViT在邻居分析中更优
- 揭示视觉变异的挑战因素,适合多模态模型研究者参考
概念的视觉表现随语义和上下文显著变化,给视觉与多模态模型带来挑战。本研究以“具体性”这一已广泛研究的词汇语义变量为案例,考察视觉表征的变异性。基于从Bing和YFCC两个数据集提取的约1000个抽象与具体概念的图像,研究目标包括:(i) 评估视觉多样性是否能可靠区分抽象与具体概念;(ii) 通过最近邻分析,评估同一概念下多幅图像间视觉特征的变异性;(iii) 通过分类与标注图像,识别导致变异的关键因素。结果表明,在区分抽象与具体概念的图像时,颜色、纹理等基础视觉特征比视觉变压器(ViT)等复杂模型提取的特征更有效;然而,ViT在最近邻分析中表现更佳,强调在通过非文本模态分析概念变量时需谨慎选择视觉特征。
原文摘要 · Abstract (English)
The visual representation of a concept varies significantly depending on its meaning and the context where it occurs; this poses multiple challenges both for vision and multimodal models. Our study focuses on concreteness, a well-researched lexical-semantic variable, using it as a case study to examine the variability in visual representations. We rely on images associated with approximately 1,000 abstract and concrete concepts extracted from two different datasets: Bing and YFCC. Our goals are: (i) evaluate whether visual diversity in the depiction of concepts can reliably distinguish between concrete and abstract concepts; (ii) analyze the variability of visual features across multiple images of the same concept through a nearest neighbor analysis; and (iii) identify challenging factors contributing to this variability by categorizing and annotating images. Our findings indicate that for classifying images of abstract versus concrete concepts, a combination of basic visual features such as color and texture is more effective than features extracted by more complex models like Vision Transformer (ViT). However, ViTs show better performances in the nearest neighbor analysis, emphasizing the need for a careful selection of visual features when analyzing conceptual variables through modalities other than text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。