仅用文本嵌入空间的分布结构,就能准确预测词语的图像性与具体性。
Uncovering Visual-Semantic Psycholinguistic Properties from the Distributional Structure of Text Embedding Space
- 基于词邻域分布的尖锐程度,提出无监督的邻里稳定性度量(NSM)
- NSM与真实评分相关性优于现有无监督方法,分类效果出色
- 无需图像或跨模态数据,适合语言模型分析与心理语言学研究
图像性(文本引发心理图像的能力)和具体性(文本的可感知性)是连接视觉与语义空间的心理语言学属性。以往计算方法通常依赖图像-标题对或多模态模型中的双模态空间。本文假设:仅从图像-标题数据集中的文本本身,即可获得准确估计这些属性的充分信号。特别地,我们提出词在语义嵌入空间中邻域的陡峭程度反映其图像性与具体性。为此,我们设计一种无监督、分布无关的度量——邻里稳定性度量(NSM),用于量化邻域峰值的尖锐性。大量实验表明,NSM与真实评分的相关性高于现有无监督方法,并在分类任务中表现优异。代码与数据已开源于GitHub。
原文摘要 · Abstract (English)
Imageability (potential of text to evoke a mental image) and concreteness (perceptibility of text) are two psycholinguistic properties that link visual and semantic spaces. It is little surprise that computational methods that estimate them do so using parallel visual and semantic spaces, such as collections of image-caption pairs or multi-modal models. In this paper, we work on the supposition that text itself in an image-caption dataset offers sufficient signals to accurately estimate these properties. We hypothesize, in particular, that the peakedness of the neighborhood of a word in the semantic embedding space reflects its degree of imageability and concreteness. We then propose an unsupervised, distribution-free measure, which we call Neighborhood Stability Measure (NSM), that quantifies the sharpness of peaks. Extensive experiments show that NSM correlates more strongly with ground-truth ratings than existing unsupervised methods, and is a strong predictor of these properties for classification. Our code and data are available on GitHub (https://github.com/Artificial-Memory-Lab/imageability).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。