arXiv:2608.04670cs.CLcs.AI2026-08被引 3

测试大模型对意大利谚语的理解,发现它们擅长补全却难选正确答案。

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark

  • 构建意大利谚语多选题基准ProverbIT,评估模型补全与选择能力。
  • 13个模型在无正确选项时准确率暴跌,顶尖推理模型也大幅退步。
  • 模型常误选字面意思,暴露对谚语深层含义理解不足。

大型语言模型(LLMs)在自然语言处理任务中表现卓越,但在理解文化嵌入的语言表达方面仍存在显著差距。本文提出ProverbIT,一个包含100道多选题的意大利语基准,用于评估模型完成意大利谚语的能力。我们测试了13个前沿模型,包括大型推理模型(LRMs)和传统LLMs,涵盖三个任务:谚语补全、带正确选项的多选、不带正确选项的多选。结果显示,几乎所有模型在补全任务中表现良好,但当移至无正确选项的多选任务时,性能急剧下降,甚至最先进的推理模型也出现显著退化。通过对两个LRMs进行链式思维分析,发现模型倾向于选择字面同义词,并在推理中提及正确谚语结尾,却未能识别其不在选项中。这表明当前模型主要依赖记忆模式,而非对文化语境下习语的深层语义理解,凸显其在隐喻语言推理上的重要局限。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have transformed computational linguistics and achieved remarkable performance across numerous natural language processing tasks, yet significant gaps persist in understanding how these systems process culturally embedded linguistic expressions. This paper introduces ProverbIT, a novel Italian benchmark comprising 100 multiple-choice questions designed to evaluate LLMs' ability to complete Italian proverbs. We assess 13 frontier models, including Large Reasoning Models (LRMs) and traditional LLMs, across three tasks: proverb completion, multiple-choice selection with correct answers, and multiple-choice selection without correct answers. Our evaluation reveals surprising results: while nearly all models demonstrate knowledge of the proverbs through successful completion tasks, performance drops dramatically when transitioning to multiple-choice formats without correct answers, with even state-of-the-art reasoning models showing substantial degradation. Through detailed Chain-of-Thought analysis of two LRMs, we uncover that models exhibit a strong bias toward selecting literal synonyms and frequently mention correct proverb endings during reasoning without successfully identifying their absence from the given options. These findings suggest that current LLMs rely heavily on memorized patterns rather than deeper semantic understanding of culturally grounded expressions, highlighting important limitations in their reasoning capabilities for figurative language comprehension.

谚语理解语言模型推理偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。