arXiv:2412.17729cs.CLcs.AI2024-12被引 3

首个大规模中文幽默理解数据集,挑战大模型真实幽默理解能力

Chumor 2.0: Towards Benchmarking Chinese Humor Understanding

  • 基于中文社交平台Ruo Zhi Ba构建超大规模中文幽默解释数据集
  • 十款大模型在该数据集上准确率仅略高于随机,远低于人类水平
  • 人工幽默解释显著优于GPT-4o和ERNIE-4-turbo,适合中文幽默研究者使用

现有幽默数据集与评估主要聚焦英语,缺乏对中文等非英语语言中文化特异性幽默的资源支持。为此,我们构建了Chumor,这是首个规模超过现有幽默数据集的中文幽默解释数据集。数据源自中文类似Reddit的平台Ruo Zhi Ba,内容涵盖智力挑战性强且具有文化特异性的笑话。我们通过直接提示与链式思维提示测试十款大语言模型,发现它们在Chumor上的准确率仅略高于随机水平,远低于人类表现。此外,分析显示人类标注的幽默解释显著优于GPT-4o和ERNIE-4-turbo生成的结果。Chumor已开源,地址为https://huggingface.co/datasets/dnaihao/Chumor,项目主页为https://dnaihao.github.io/Chumor-dataset/,排行榜位于https://huggingface.co/spaces/dnaihao/Chumor,代码库为https://github.com/dnaihao/Chumor-dataset。

原文摘要 · Abstract (English)

Existing humor datasets and evaluations predominantly focus on English, leaving limited resources for culturally nuanced humor in non-English languages like Chinese. To address this gap, we construct Chumor, the first Chinese humor explanation dataset that exceeds the size of existing humor datasets. Chumor is sourced from Ruo Zhi Ba, a Chinese Reddit-like platform known for sharing intellectually challenging and culturally specific jokes. We test ten LLMs through direct and chain-of-thought prompting, revealing that Chumor poses significant challenges to existing LLMs, with their accuracy slightly above random and far below human. In addition, our analysis highlights that human-annotated humor explanations are significantly better than those generated by GPT-4o and ERNIE-4-turbo. We release Chumor at https://huggingface.co/datasets/dnaihao/Chumor, our project page is at https://dnaihao.github.io/Chumor-dataset/, our leaderboard is at https://huggingface.co/spaces/dnaihao/Chumor, and our codebase is at https://github.com/dnaihao/Chumor-dataset.

中文幽默大模型评测数据集发布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。