大模型在精确量词判断上比模糊量词更像人类,挑战了分布语义的普遍认知。
Are LLMs Models of Distributional Semantics? A Case Study on Quantifiers
- 用精确与模糊量词对比测试大模型语义理解能力
- 多数模型在精确量词上表现优于模糊量词,接近人类判断
- 研究质疑分布语义理论对模型能力的预设,适合语言计算研究者
分布语义认为词语意义源于其在自然语言中的使用模式。语言模型常被视为分布语义的实现,因其被优化以捕捉语言的统计特征。通常认为,分布语义模型应擅长处理基于语言惯例的模糊意义,但在真值条件推理和符号处理上表现不佳。本文通过精确量词(如“超过一半”)与模糊量词(如“许多”)的案例研究检验此观点。结果出人意料:在多种类型的大模型中,模型在精确量词上的判断更接近人类,而对模糊量词的表现则较差。这一发现要求重新审视分布语义模型的本质及其可捕获的能力边界。
原文摘要 · Abstract (English)
Distributional semantics is the linguistic theory that a word's meaning can be derived from its distribution in natural language (i.e., its use). Language models are commonly viewed as an implementation of distributional semantics, as they are optimized to capture the statistical features of natural language. It is often argued that distributional semantics models should excel at capturing graded/vague meaning based on linguistic conventions, but struggle with truth-conditional reasoning and symbolic processing. We evaluate this claim with a case study on vague (e.g. "many") and exact (e.g. "more than half") quantifiers. Contrary to expectations, we find that, across a broad range of models of various types, LLMs align more closely with human judgements on exact quantifiers versus vague ones. These findings call for a re-evaluation of the assumptions underpinning what distributional semantics models are, as well as what they can capture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。