大模型能否跨类型理解笑话?实验发现能,但效果有限。
One Joke to Rule them All? On the (Im)possibility of Generalizing Humor
- 在多个笑话数据集上训练大模型,测试其跨类型迁移能力。
- 模型在未见过的笑话类型上最高达75%准确率,多源训练提升迁移性。
- 老爸笑话最助迁移,但难被其他类型学习,适合研究幽默通用性者。
幽默是一种广泛而复杂的沟通形式,对机器仍具挑战。现有研究多聚焦特定幽默类型,本文探究掌握一种或多种幽默任务是否能迁移到全新未见类型,即这种碎片化是否不可避免。随着网络与社交媒体中新笑点(如梗图、反幽默、AI翻车)不断涌现,大语言模型若要适应,必须具备跨类型泛化能力。为此,我们在四个不同幽默任务的数据集上开展迁移学习实验,训练时使用1至3个数据集,测试在新任务上的表现。结果表明,模型具备一定迁移能力,在未见数据集上最高可达75%准确率;多源训练可使迁移性能提升1.88-4.05%,且不影响原域表现。进一步分析显示各类幽默间存在关联,其中老爸笑话虽难被迁移,却成为最佳迁移促进者。相关数据与代码已开源。
原文摘要 · Abstract (English)
Humor is a broad and complex form of communication that remains challenging for machines. Despite its broadness, most existing research on computational humor traditionally focused on modeling a specific type of humor. In this work, we wish to understand whether competence on one or more specific humor tasks confers any ability to transfer to novel, unseen types; in other words, is this fragmentation inevitable? This question is especially timely as new humor types continuously emerge in online and social media contexts (e.g., memes, anti-humor, AI fails). If Large Language Models (LLMs) are to keep up with this evolving landscape, they must be able to generalize across humor types by capturing deeper, transferable mechanisms. To investigate this, we conduct a series of transfer learning experiments across four datasets, representing different humor tasks. We train LLMs under varied diversity settings (1-3 datasets in training, testing on a novel task). Experiments reveal that models are capable of some transfer, and can reach up to 75% accuracy on unseen datasets; training on diverse sources improves transferability (1.88-4.05%) with minimal-to-no drop in in-domain performance. Further analysis suggests relations between humor types, with Dad Jokes surprisingly emerging as the best enabler of transfer (but is difficult to transfer to). We release data and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。