评测大模型幽默生成与判断能力,发现其缺共情、重新奇。
Assessing the Capabilities of LLMs in Humor:A Multi-dimensional Analysis of Oogiri Generation and Evaluation
- 用日本即兴喜剧Oogiri多维度评测大模型幽默能力
- 模型生成水平达人类中低段,但共情力严重不足
- 适合研究情感智能对话系统的人参考
计算幽默是推动先进自然语言处理应用(如复杂对话系统)的前沿领域。以往研究多依赖单一维度评价(如是否好笑),本文提出需多维理解幽默,并以日本即兴喜剧Oogiri为视角系统评估大模型。我们扩充了现有Oogiri数据集,新增来源并引入大模型生成响应,再由人工标注六维评分:新颖性、清晰度、相关性、智慧性、共情力、整体好笑度。基于此数据集,评估了当前顶尖大模型在生成和评估两任务中的表现。结果表明,模型生成水平介于人类低至中等之间,但在共情维度显著缺失;相关性分析显示,模型偏重新颖性,而人类更看重共情。我们公开该标注语料库,助力更富情感智能对话系统的发展。
原文摘要 · Abstract (English)
Computational humor is a frontier for creating advanced and engaging natural language processing (NLP) applications, such as sophisticated dialogue systems. While previous studies have benchmarked the humor capabilities of Large Language Models (LLMs), they have often relied on single-dimensional evaluations, such as judging whether something is simply ``funny.'' This paper argues that a multifaceted understanding of humor is necessary and addresses this gap by systematically evaluating LLMs through the lens of Oogiri, a form of Japanese improvisational comedy games. To achieve this, we expanded upon existing Oogiri datasets with data from new sources and then augmented the collection with Oogiri responses generated by LLMs. We then manually annotated this expanded collection with 5-point absolute ratings across six dimensions: Novelty, Clarity, Relevance, Intelligence, Empathy, and Overall Funniness. Using this dataset, we assessed the capabilities of state-of-the-art LLMs on two core tasks: their ability to generate creative Oogiri responses and their ability to evaluate the funniness of responses using a six-dimensional evaluation. Our results show that while LLMs can generate responses at a level between low- and mid-tier human performance, they exhibit a notable lack of Empathy. This deficit in Empathy helps explain their failure to replicate human humor assessment. Correlation analyses of human and model evaluation data further reveal a fundamental divergence in evaluation criteria: LLMs prioritize Novelty, whereas humans prioritize Empathy. We release our annotated corpus to the community to pave the way for the development of more emotionally intelligent and sophisticated conversational agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。