arXiv:2601.09041cs.CLcs.AI2026-01被引 1

大模型能模仿人类表层判断,但理解隐喻和俚语时仍差一截。

Can LLMs interpret figurative language as humans do?: surface-level vs representational similarity

  • 用40个问题测评240句对话,对比人与大模型的判断差异
  • 大模型在隐喻和年轻俚语上与人类代表层面差距明显
  • GPT-4最接近人类,但对讽刺、习语等社会语用表达仍吃力

大型语言模型生成的判断看似接近人类。然而,这些模型在解读具有社会语境的隐喻性语言方面与人类的一致性仍不明确。为此,研究者邀请人类参与者,并测试了四种不同规模的指令微调大模型(GPT-4、Gemma-2-9B、Llama-3.2、Mistral-7B),对240个基于对话的句子进行评估,涵盖六种语言特征:惯例性、讽刺、搞笑、情感、习语性和俚语性。每句话配以40个解释性问题,人类与模型均使用10分李克特量表评分。结果表明,人类与模型在表面层面上一致,但在表征层面显著分歧,尤其在处理涉及习语和青年俚语的隐喻句子时。尽管GPT-4最接近人类的表征模式,所有模型在理解依赖语境和社会语用的表达(如讽刺、俚语、习语性)上仍表现不佳。

原文摘要 · Abstract (English)

Large language models generate judgments that resemble those of humans. Yet the extent to which these models align with human judgments in interpreting figurative and socially grounded language remains uncertain. To investigate this, human participants and four instruction-tuned LLMs of different sizes (GPT-4, Gemma-2-9B, Llama-3.2, and Mistral-7B) rated 240 dialogue-based sentences representing six linguistic traits: conventionality, sarcasm, funny, emotional, idiomacy, and slang. Each of the 240 sentences was paired with 40 interpretive questions, and both humans and LLMs rated these sentences on a 10-point Likert scale. Results indicated that humans and LLMs aligned at the surface level with humans, but diverged significantly at the representational level, especially in interpreting figurative sentences involving idioms and Gen Z slang. GPT-4 most closely approximates human representational patterns, while all models struggle with context-dependent and socio-pragmatic expressions like sarcasm, slang, and idiomacy.

大模型理解隐喻识别语用分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。