用图像数据量化词语在不同语言中的语义实质,发现名词最具体、语法词也含信息。
A Grounded Typology of Word Classes
- 以图像为语义基准,通过多语言多模态模型计算词语的语义充实度。
- 发现名词>形容词>动词的普遍语义充实度层级,且语法词也有内容信息。
- 适用于研究语言共性、跨语言语义差异或心理语言学的定量分析。
我们提出一种基于感知模态(如图像)的语言类型学意义量化方法,将图像作为无语言依赖的语义表示,从而跨语言量化图像与描述之间的功能-形式关系。受信息论启发,定义了‘具象度’这一可实证的上下文语义充实度指标(以意外度差异形式表达),并利用多语言多模态语言模型进行计算。作为概念验证,我们将其应用于词类类型学研究:结果揭示出功能词(语法词)与词汇词(实义词)之间语义充实度的不对称性,挑战了‘功能词不传递语义’的传统观点;同时发现普遍存在的具象度层级(如名词 > 形容词 > 动词),且该指标部分与英语心理语言学具体性量表相关。我们发布了30种语言的具象度评分数据集。结果表明,具象类型学方法可为语言语义功能提供可量化的证据。
原文摘要 · Abstract (English)
We propose a grounded approach to meaning in language typology. We treat data from perceptual modalities, such as images, as a language-agnostic representation of meaning. Hence, we can quantify the function--form relationship between images and captions across languages. Inspired by information theory, we define "groundedness", an empirical measure of contextual semantic contentfulness (formulated as a difference in surprisal) which can be computed with multilingual multimodal language models. As a proof of concept, we apply this measure to the typology of word classes. Our measure captures the contentfulness asymmetry between functional (grammatical) and lexical (content) classes across languages, but contradicts the view that functional classes do not convey content. Moreover, we find universal trends in the hierarchy of groundedness (e.g., nouns > adjectives > verbs), and show that our measure partly correlates with psycholinguistic concreteness norms in English. We release a dataset of groundedness scores for 30 languages. Our results suggest that the grounded typology approach can provide quantitative evidence about semantic function in language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。