arXiv:2507.15357cs.CLcs.AI2025-07ACL被引 14

大模型理解隐喻靠表面特征,而非真正懂意思。

Metaphor and Large Language Models: When Surface Features Matter More than Deep Understanding

  • 用多数据集测试大模型隐喻理解能力
  • 表现受词汇重叠和句子长度影响更大
  • 适合研究语言模型真实理解力的学者

本文对大语言模型(LLMs)在多种数据集、任务和提示配置下处理隐喻的能力进行了全面评估。尽管隐喻理解在自然语言处理中备受关注,但以往研究多局限于单一数据集、特定任务设置,且常使用通过词替换生成的人工数据。本研究克服这些局限,采用多个公开可用的数据集,包含推理与隐喻标注,聚焦自然语言推理(NLI)和问答(QA)任务,开展广泛实验。结果表明,大模型的表现更受词汇重叠、句子长度等表面特征影响,而非隐喻内容本身,说明其所谓“涌现”的隐喻理解能力,实为表面特征、上下文学习与语言知识共同作用的结果。该研究揭示了当前大模型在处理比喻语言时的能力边界,强调需建立更真实的评估框架。数据与代码已公开。

原文摘要 · Abstract (English)

This paper presents a comprehensive evaluation of the capabilities of Large Language Models (LLMs) in metaphor interpretation across multiple datasets, tasks, and prompt configurations. Although metaphor processing has gained significant attention in Natural Language Processing (NLP), previous research has been limited to single-dataset evaluations and specific task settings, often using artificially constructed data through lexical replacement. We address these limitations by conducting extensive experiments using diverse publicly available datasets with inference and metaphor annotations, focusing on Natural Language Inference (NLI) and Question Answering (QA) tasks. The results indicate that LLMs' performance is more influenced by features like lexical overlap and sentence length than by metaphorical content, demonstrating that any alleged emergent abilities of LLMs to understand metaphorical language are the result of a combination of surface-level features, in-context learning, and linguistic knowledge. This work provides critical insights into the current capabilities and limitations of LLMs in processing figurative language, highlighting the need for more realistic evaluation frameworks in metaphor interpretation tasks. Data and code are publicly available.

隐喻理解大模型评测语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。