arXiv:2502.05704cs.CLcs.AI2025-02NAACL被引 5

用分类混淆度衡量词语语义相似性,更贴合语言的动态变化。

Rethinking Word Similarity: Semantic Similarity through Classification Confusion

  • 基于上下文嵌入训练分类器,用误判概率衡量词语相似性
  • 在多个数据集上表现接近传统方法,且可动态选择特征
  • 适合研究词语意义随时间演变的文化分析任务

词语相似性在社会科学研究和文化分析中具有广泛应用,如测量语义随时间的变化或解析争议性术语。然而,基于词向量余弦相似性的传统方法难以捕捉语义相似性的上下文依赖、非对称性和多义性。本文提出一种新度量方式——词混淆度(Word Confusion),将语义相似性重新定义为基于特征的分类混淆。该方法受Tversky启发,动态选择相似性特征:训练分类器将上下文嵌入映射到词身份,并以误判为干扰词c的概率(而非正确目标词t)作为c与t的相似性度量。潜在干扰词集合即为所选特征。实验表明,该方法在MEN、WordSim353和SimLex三个数据集上与人类判断的匹配度可与余弦相似性媲美,且能利用预设特征进行相似性测量。通过应用于18世纪法语词'révolution'的意义演变研究,验证了其在动态特征下的适用性。本工作旨在推动计算社会科学与文化分析等领域对语言多面性与动态性的更精准建模。

原文摘要 · Abstract (English)

Word similarity has many applications to social science and cultural analytics tasks like measuring meaning change over time and making sense of contested terms. Yet traditional similarity methods based on cosine similarity between word embeddings cannot capture the context-dependent, asymmetrical, polysemous nature of semantic similarity. We propose a new measure of similarity, Word Confusion, that reframes semantic similarity in terms of feature-based classification confusion. Word Confusion is inspired by Tversky's suggestion that similarity features be chosen dynamically. Here we train a classifier to map contextual embeddings to word identities and use the classifier confusion (the probability of choosing a confounding word c instead of the correct target word t) as a measure of the similarity of c and t. The set of potential confounding words acts as the chosen features. Our method is comparable to cosine similarity in matching human similarity judgments across several datasets (MEN, WirdSim353, and SimLex), and can measure similarity using predetermined features of interest. We demonstrate our model's ability to make use of dynamic features by applying it to test a hypothesis about changes in the 18th C. meaning of the French word "revolution" from popular to state action during the French Revolution. We hope this reimagining of semantic similarity will inspire the development of new tools that better capture the multi-faceted and dynamic nature of language, advancing the fields of computational social science and cultural analytics and beyond.

语义相似性动态语义分类混淆文化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。