LLM时代下,词义归纳仍无解,现有方法难超基础聚类策略。
In the LLM era, Word Sense Induction remains unsolved
- 基于语料原生多义性与频次分布构建评估集,更贴近真实场景。
- 无论何种方法,均无法超越‘每词一簇’的简单启发式,最优仅提升3.3%。
- 利用维基词典增强数据有效,但大模型在词义归纳中表现不佳。
在缺乏词义标注数据的情况下,词义归纳(WSI)是低资源或特定领域中词义消歧的有力替代方案。本文指出当前WSI评估中的方法论问题,提出基于SemCor衍生数据集的评估方案,保留原始语料的多义性和频率分布。我们评估了预训练词向量与聚类算法在不同词性下的表现,并提出并测试了一种基于大语言模型的英文WSI方法。研究涵盖多种数据增强来源(LLM生成、语料库、词典)、半监督场景(使用Wiktionary进行数据增强、必须共现约束、每词簇数设定)。结果表明:(i)不同词性下性能差异显著;(ii)大模型在该任务中表现不佳;(iii)数据增强有效;(iv)利用维基词典确实有帮助。所提方法在测试集上优于先前最强系统3.3%。词义归纳仍未解决,亟需更清晰的词典构建与对大模型词汇语义能力的重新审视。
原文摘要 · Abstract (English)
In the absence of sense-annotated data, word sense induction (WSI) is a compelling alternative to word sense disambiguation, particularly in low-resource or domain-specific settings. In this paper, we emphasize methodological problems in current WSI evaluation. We propose an evaluation on a SemCor-derived dataset, respecting the original corpus polysemy and frequency distributions. We assess pre-trained embeddings and clustering algorithms across parts of speech, and propose and evaluate an LLM-based WSI method for English. We evaluate data augmentation sources (LLM-generated, corpus and lexicon), and semi-supervised scenarios using Wiktionary for data augmentation, must-link constraints, number of clusters per lemma. We find that no unsupervised method (whether ours or previous) surpasses the strong "one cluster per lemma" heuristic (1cpl). We also show that (i) results and best systems may vary across POS, (ii) LLMs have troubles performing this task, (iii) data augmentation is beneficial and (iv) capitalizing on Wiktionary does help. It surpasses previous SOTA system on our test set by 3.3\%. WSI is not solved, and calls for a better articulation of lexicons and LLMs' lexical semantics capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。