arXiv:2508.04494cs.CL2025-08Conference of the …

通过概念对齐提升词语语义表示,兼顾同词多义与跨词语义区分。

CALE : Concept-Aligned Embeddings for Both Within-Lemma and Inter-Lemma Sense Differentiation

  • 引入概念区分任务,扩展传统同词多义研究到跨词语义比较。
  • 基于SemCor构建数据集,微调模型后在多个语义任务上表现最优。
  • 改善嵌入空间结构,使语义分布更符合人类认知逻辑。

词汇语义关注词语在不同语境下的多重含义,以及不同词语意义间的语义关系。上下文语言模型能提供情境敏感的表示,是研究词汇意义的重要工具。近期工作如XL-LEXEME利用词语上下文任务微调模型以获得更准确的语义表示,但该任务仅比较同一词素的不同用法,限制了信息覆盖范围。本文提出概念区分(Concept Differentiation)作为扩展,涵盖跨词语义场景。我们基于SemCor数据构建了相应数据集,并在该数据集上微调多种表示模型,命名为概念对齐嵌入(CALE)。通过在多个词汇语义任务上测试我们的模型及其他模型,结果表明,所提模型提供了高效、多功能的词汇意义表示,在实验中达到最佳性能。此外,我们还发现CALE的微调显著改变了嵌入的空间组织结构。

原文摘要 · Abstract (English)

Lexical semantics is concerned with both the multiple senses a word can adopt in different contexts, and the semantic relations that exist between meanings of different words. To investigate them, Contextualized Language Models are a valuable tool that provides context-sensitive representations that can be used to investigate lexical meaning. Recent works like XL-LEXEME have leveraged the task of Word-in-Context to fine-tune them to get more semantically accurate representations, but Word-in-Context only compares occurrences of the same lemma, limiting the range of captured information. In this paper, we propose an extension, Concept Differentiation, to include inter-words scenarios. We provide a dataset for this task, derived from SemCor data. Then we fine-tune several representation models on this dataset. We call these models Concept-Aligned Embeddings (CALE). By challenging our models and other models on various lexical semantic tasks, we demonstrate that the proposed models provide efficient multi-purpose representations of lexical meaning that reach best performances in our experiments. We also show that CALE's fine-tuning brings valuable changes to the spatial organization of embeddings.

语义表示词义消歧嵌入对齐上下文模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。