arXiv:2509.12158cs.CLcs.AI2025-09EMNLP被引 5

大模型看懂双关语只是表面功夫,一改细节就懵了。

Pun Unintended: LLMs and the Illusion of Humor Understanding

  • 重构双关语测试集,揭示模型依赖表面特征
  • 微调语义或发音即导致模型误判率飙升
  • 适合研究模型幽默理解局限的读者

双关语是一种利用多义性和语音相似性的幽默语言游戏。尽管大模型在识别双关语方面展现出潜力,但本文表明其理解往往停留在浅层,缺乏人类解读所需的细腻把握。通过系统分析并重新构建现有双关语评测基准,我们证明了双关语的细微变化即可误导大模型。主要贡献包括全面且细致的双关语检测评测基准、对近期大模型的人类评估,以及对模型处理双关语时鲁棒性挑战的分析。

原文摘要 · Abstract (English)

Puns are a form of humorous wordplay that exploits polysemy and phonetic similarity. While LLMs have shown promise in detecting puns, we show in this paper that their understanding often remains shallow, lacking the nuanced grasp typical of human interpretation. By systematically analyzing and reformulating existing pun benchmarks, we demonstrate how subtle changes in puns are sufficient to mislead LLMs. Our contributions include comprehensive and nuanced pun detection benchmarks, human evaluation across recent LLMs, and an analysis of the robustness challenges these models face in processing puns.

双关语大模型幽默理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。