构建幽默理解基准,让大模型学懂日本冷笑话的笑点
Oogiri-Master: Benchmarking Humor Understanding via Oogiri
- 用100个独立评委对每条梗做匿名评分,避免从众效应
- 发现长度、歧义和意外反转是决定笑点的关键语言特征
- 首次实现模型与人类在冷笑话生成上接近同水平
幽默是检验大语言模型类人创造力的重要场景。本文以日本创意回应游戏Oogiri为研究对象,探究何种回应能被人类认为好笑。现有数据集存在每题候选回复少、评分时暴露流行度信号、缺乏客观可比的幽默度指标等问题。为此,我们提出Oogiri-Master基准和Oogiri-Corpus数据集:每条提示对应约100条多样回复,由约100名独立人类评委匿名评分,有效降低流行度偏差并实现稳健聚合。基于该数据集,我们定量分析了文本长度、歧义性及意外性化解等语言因素与幽默感的关系,并建立可预测人类判断的客观指标。随后在Oogiri-Master上评估多种大模型与人类基线,结果显示顶尖模型已接近人类表现,且引入洞察提示可进一步提升性能。本工作为幽默理解的评估与进步提供了系统化基础。
原文摘要 · Abstract (English)
Humor is a salient testbed for human-like creative thinking in large language models (LLMs). We study humor using the Japanese creative response game Oogiri, in which participants produce witty responses to a given prompt, and ask the following research question: What makes such responses funny to humans? Previous work has offered only limited reliable means to answer this question. Existing datasets contain few candidate responses per prompt, expose popularity signals during ratings, and lack objective and comparable metrics for funniness. Thus, we introduce Oogiri-Master and Oogiri-Corpus, which are a benchmark and dataset designed to enable rigorous evaluation of humor understanding in LLMs. Each prompt is paired with approximately 100 diverse candidate responses, and funniness is rated independently by approximately 100 human judges without access to others' ratings, reducing popularity bias and enabling robust aggregation. Using Oogiri-Corpus, we conduct a quantitative analysis of the linguistic factors associated with funniness, such as text length, ambiguity, and incongruity resolution, and derive objective metrics for predicting human judgments. Subsequently, we benchmark a range of LLMs and human baselines in Oogiri-Master, demonstrating that state-of-the-art models approach human performance and that insight-augmented prompting improves the model performance. Our results provide a principled basis for evaluating and advancing humor understanding in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。