测试大模型对英文'just'的语用细微差别理解能力,发现仍存明显短板。
Is It JUST Semantics? A Case Study of Discourse Particle Understanding in LLMs
- 用专家标注的语料测试模型对'just'多义性的分辨能力
- 模型能区分大类但难把握精细语义差异
- 适合关注语言理解局限的研究者参考
话语标记词是微妙塑造文本意义的关键元素。这些常具多重功能的词语,会产生丰富且差异显著的语义/语用效果,以英语中的'just'为例(如排他性、时间性、强调等)。本研究利用语言学家精心构建并标注的数据,考察大语言模型对英语'just'细粒度语义的理解能力。结果表明,尽管模型具备一定区分宽泛类别能力,却难以完整捕捉更细微的语义差别,揭示了其在话语标记词理解上的显著不足。
原文摘要 · Abstract (English)
Discourse particles are crucial elements that subtly shape the meaning of text. These words, often polyfunctional, give rise to nuanced and often quite disparate semantic/discourse effects, as exemplified by the diverse uses of the particle "just" (e.g., exclusive, temporal, emphatic). This work investigates the capacity of LLMs to distinguish the fine-grained senses of English "just", a well-studied example in formal semantics, using data meticulously created and labeled by expert linguists. Our findings reveal that while LLMs exhibit some ability to differentiate between broader categories, they struggle to fully capture more subtle nuances, highlighting a gap in their understanding of discourse particles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。