用语义网络分析词义,让模型更懂上下文。
SANDWiCH: Semantical Analysis of Neighbours for Disambiguating Words in Context ad Hoc
- 基于贝宝网构建语义网络,用群代数优化词义判别。
- 多语言词义消歧全任务刷新纪录,低资源语言表现突出。
- 参数量减少72%,效率高适合实际部署。
过去两年生成式对话大模型的兴起推动了近似人类对话与推理体验系统的快速发展。然而,近期研究表明,这些模型的语言理解能力仍有限,远未达到人类水平,尤其在把握词语上下文含义方面,而这是推理的核心。本文提出一种简单且计算高效的多语言词义消歧(WSD)框架。我们将WSD任务重新定义为基于贝宝网(BabelNet)经群代数优化后的语义网络上的聚类判别分析。在多个WSD基准上验证方法,对所有语言和任务均达到新最佳性能,且在词性分类的独立评估中同样领先。值得注意的是,该模型显著超越现有方法,即使在低资源语言中也表现优异,同时参数量减少72%。
原文摘要 · Abstract (English)
The rise of generative chat-based Large Language Models (LLMs) over the past two years has spurred a race to develop systems that promise near-human conversational and reasoning experiences. However, recent studies indicate that the language understanding offered by these models remains limited and far from human-like performance, particularly in grasping the contextual meanings of words, an essential aspect of reasoning. In this paper, we present a simple yet computationally efficient framework for multilingual Word Sense Disambiguation (WSD). Our approach reframes the WSD task as a cluster discrimination analysis over a semantic network refined from BabelNet using group algebra. We validate our methodology across multiple WSD benchmarks, achieving a new state of the art for all languages and tasks, as well as in individual assessments by part of speech. Notably, our model significantly surpasses the performance of current alternatives, even in low-resource languages, while reducing the parameter count by 72%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。