测试大模型是否真懂词语在语境中的意思,发现表现接近专业系统。
Do Large Language Models Understand Word Senses?
- 用多任务评估大模型的词义消歧与生成能力。
- 生成任务中词义理解准确率达98%,自由解释表现最优。
- 闭源和开源模型均表现出强鲁棒性,适合语言理解研究。
理解词语在语境中的含义是大语言模型(LLM)的基本能力。尽管已有大量评估工作,但大模型是否真正掌握词义仍缺乏深入探究。本文通过两项评估填补这一空白:一是对比指令微调的大模型与专用词义消歧(WSD)系统在WSD任务上的表现;二是考察两个顶尖开源与闭源模型在三种生成任务中的词义理解能力:定义生成、自由解释和示例生成。结果表明,在WSD任务中,GPT-4o和DeepSeek-V3等领先模型的表现与专门设计的WSD系统相当,且在不同领域和难度下更具鲁棒性。在生成任务中,模型对上下文词义的理解准确率最高达98%,其中自由解释任务表现最佳,最契合其生成能力。
原文摘要 · Abstract (English)
Understanding the meaning of words in context is a fundamental capability for Large Language Models (LLMs). Despite extensive evaluation efforts, the extent to which LLMs show evidence that they truly grasp word senses remains underexplored. In this paper, we address this gap by evaluating both i) the Word Sense Disambiguation (WSD) capabilities of instruction-tuned LLMs, comparing their performance to state-of-the-art systems specifically designed for the task, and ii) the ability of two top-performing open- and closed-source LLMs to understand word senses in three generative settings: definition generation, free-form explanation, and example generation. Notably, we find that, in the WSD task, leading models such as GPT-4o and DeepSeek-V3 achieve performance on par with specialized WSD systems, while also demonstrating greater robustness across domains and levels of difficulty. In the generation tasks, results reveal that LLMs can explain the meaning of words in context up to 98\% accuracy, with the highest performance observed in the free-form explanation task, which best aligns with their generative capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。