arXiv:2606.18922cs.CLcs.AI2026-06

测试大模型在讽刺与否定共现时的理解能力,发现提示方式影响关键表现。

As Easy as Rocket Science: Assessing the Ability of Large Language Models to Interpret Negation in Figurative Language

论文配图:As Easy as Rocket Science: Assessing the Ability of Large Language Models to Interpret Negation in Figurative Language
图 1 · 摘自论文原文
  • 构建新标注数据集,研究否定与隐喻并存文本的解析机制。
  • 模型整体表现差,不同否定类型下准确率差异显著。
  • 提示风格对结果影响大,适合评估通用语言理解能力的研究者参考。

隐喻和否定是当前语言模型面临的主要挑战,但二者在书面和口语中广泛存在。大语言模型(LLMs)常用于无需特定微调的日常场景,因此了解其对包含否定和隐喻的文本的正确理解能力至关重要。为此,我们对现有隐喻数据集进行了新的标注,并测试了多种语言模型的表现。结果发现,否定与隐喻的结合构成特殊难题,且整体性能及不同否定类型下的表现高度依赖于提示风格。

原文摘要 · Abstract (English)

Figurative language and negation are two areas that challenge current language models, however, both are widely used throughout written and spoken language. Large language models (LLMs) are also widely used in everyday contexts where they cannot necessarily be tuned for a specific dataset. It is therefore essential to understand the ability of LLMs to correctly interpret text that includes both negation and figurative language. To investigate this, we develop a set of new annotations to an existing dataset of figurative language, and test a range of language models on the dataset. We find that the combination of negation and figurativeness can present a particular challenge, and that performance overall and across different negation types is particularly dependent on the prompt style used.

语言模型隐喻理解否定处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。