arXiv:2504.09049cs.CL2025-04中稿 · NAACL被引 3

评测大模型识别单口喜剧笑点能力,发现顶尖模型准确率仅51%。

From Punchlines to Predictions: A Metric to Assess LLM Performance in Identifying Humor in Stand-Up Comedy

  • 设计模块化幽默检测指标,融合字符串匹配、语义嵌入与子空间相似度
  • 顶尖模型在喜剧段子识别中最高准确率51%,优于人类的41%
  • 揭示幽默主观性,适合对幽默理解与评估感兴趣的AI研究者

喜剧深刻反映时代背景,是人际互动的重要组成部分。随着大语言模型(LLMs)广泛应用,幽默与人工智能的交汇已不容忽视。自然的人机交互进展,取决于模型对幽默的理解能力。本研究评估模型从单口喜剧转录文本中准确识别幽默金句的能力。单口喜剧独特的叙事结构使其成为提升幽默理解自然性的理想数据集。我们提出一种新型幽默检测指标,可针对不同提示(prompt)评估LLM提取笑点的能力,具备模糊字符串匹配、句子嵌入和子空间相似度三种评分方式。模型结果与人类评估者对比显示:无论提示工程如何优化,领先模型ChatGPT、Claude和DeepSeek的最高幽默识别准确率为51%,高于人类的41%。人类与模型间的一致性分析揭示了幽默的主观性及从现场表演转录中提取笑点的复杂性。代码已开源于https://github.com/swaggirl9000/humor。

原文摘要 · Abstract (English)

Comedy serves as a profound reflection of the times we live in and is a staple element of human interactions. In light of the widespread adoption of Large Language Models (LLMs), the intersection of humor and AI has become no laughing matter. Advancements in the naturalness of human-computer interaction correlates with improvements in AI systems' abilities to understand humor. In this study, we assess the ability of models in accurately identifying humorous quotes from a stand-up comedy transcript. Stand-up comedy's unique comedic narratives make it an ideal dataset to improve the overall naturalness of comedic understanding. We propose a novel humor detection metric designed to evaluate LLMs amongst various prompts on their capability to extract humorous punchlines. The metric has a modular structure that offers three different scoring methods - fuzzy string matching, sentence embedding, and subspace similarity - to provide an overarching assessment of a model's performance. The model's results are compared against those of human evaluators on the same task. Our metric reveals that regardless of prompt engineering, leading models, ChatGPT, Claude, and DeepSeek, achieve scores of at most 51% in humor detection. Notably, this performance surpasses that of humans who achieve a score of 41%. The analysis of human evaluators and LLMs reveals variability in agreement, highlighting the subjectivity inherent in humor and the complexities involved in extracting humorous quotes from live performance transcripts. Code available at https://github.com/swaggirl9000/humor.

幽默识别大模型评估喜剧理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。