arXiv:2502.09743cs.CL2025-02ACL被引 3

用部分共指关系提升概念嵌入效果,让跨语言数据更可用

Partial Colexifications Improve Concept Embeddings

  • 基于部分共指网络构建概念嵌入,突破传统词级限制
  • 在语义相似度、语义演变等任务中表现更优
  • 适合低资源语言与跨语言研究者使用

尽管词嵌入已彻底改变自然语言处理领域,但概念嵌入仍受关注较少。然而,稠密且有意义的概念表示对计算语言学中的多项任务(尤其是跨语言或低资源语言的稀疏数据)具有潜在价值。现有方法基于自动构建的共指网络进行概念嵌入,但局限于词级,忽略了仅在部分词形上成立的词汇关系。本文基于近期提出的部分共指推断方法,展示了如何有效提升概念嵌入质量。所学嵌入在词汇相似度评分、语义演变实例和词关联数据上均表现更优。结果表明,该方法能准确捕捉并表达概念间的多种语义关系。

原文摘要 · Abstract (English)

While the embedding of words has revolutionized the field of Natural Language Processing, the embedding of concepts has received much less attention so far. A dense and meaningful representation of concepts, however, could prove useful for several tasks in computational linguistics, especially those involving cross-linguistic data or sparse data from low resource languages. First methods that have been proposed so far embed concepts from automatically constructed colexification networks. While these approaches depart from automatically inferred polysemies, attested across a larger number of languages, they are restricted to the word level, ignoring lexical relations that would only hold for parts of the words in a given language. Building on recently introduced methods for the inference of partial colexifications, we show how they can be used to improve concept embeddings in meaningful ways. The learned embeddings are evaluated against lexical similarity ratings, recorded instances of semantic shift, and word association data. We show that in all evaluation tasks, the inclusion of partial colexifications lead to improved concept representations and better results. Our results further show that the learned embeddings are able to capture and represent different semantic relationships between concepts.

概念嵌入共指关系低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。