对比不同笑话类型,发现大模型理解幽默能力存在明显短板
Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokes
- 构建包含600个笑话的多类型数据集,涵盖双关、网络梗等
- 测试多个大模型在零样本下解释不同笑话的能力,效果普遍不佳
- 揭示当前研究过度关注简单双关,忽视复杂现实梗的理解
幽默作为一种复杂的语言形式,源于生活的多个方面。尽管现有计算幽默研究几乎仅聚焦于简短的双关笑话,本文探究大语言模型(LLMs)理解幽默的能力是否依赖于笑话的具体形式。我们比较了模型对从简单双关到需掌握现实世界知识的复杂时事笑话的解释能力。为此,我们构建了一个包含600个笑话的数据集,涵盖异形/同音双关、当代网络幽默和时事笑话,并人工撰写高质量解释。利用该数据集,我们评估了多种LLMs在零样本下的解释能力,识别出幽默解释任务中的关键研究空白。结果表明,所有测试模型(包括推理增强模型)均无法可靠生成所有类型笑话的充分解释,进一步凸显现有研究对过于简单的笑话形式的片面关注。
原文摘要 · Abstract (English)
Humour, as a complex language form, is derived from myriad aspects of life. Whilst existing work on computational humour has focussed almost exclusively on short pun-based jokes, we investigate whether the ability of Large Language Models (LLMs) to explain humour depends on the particular form. We compare models' joke explanation abilities from simple puns to complex topical humour that requires esoteric knowledge of real-world entities and events. To this end, we curate a dataset of 600 jokes across 4 joke types and manually write high-quality explanations. These jokes include heterographic and homographic puns, contemporary internet humour, and topical jokes. Using this dataset, we compare the zero-shot abilities of a range of LLMs to accurately and comprehensively explain jokes of different types, identifying key research gaps in the task of humour explanation. We find that none of the tested models (including reasoning models) are capable of reliably generating adequate explanations of all joke types, further highlighting the narrow focus of most existing works on overly simple joke forms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。