首个专攻中文幽默生成与理解的大模型,让AI也能讲好中国笑话。
CFunModel: A "Funny" Language Model Capable of Chinese Humor Generation and Processing
- 基于2万条贴吧笑话构建16万条中文幽默数据集,覆盖多种笑点类型。
- 在相声回应选择、幽默识别、笑话生成任务中超越主流大模型。
- 适合对中文语用、幽默计算或内容生成感兴趣的开发者与研究者。
幽默在日常语言交流中至关重要。随着大语言模型(LLMs)的快速发展,自然语言处理在各类文本的理解与生成上取得显著进展。然而,现有大模型在中文幽默的生成与理解方面表现欠佳。为此,本文构建了首个全面的中文幽默数据集——中文趣集(CFunSet),整合现有中文幽默数据集,并从知名笑话分享平台贴吧-段子吧收集超过20,000条笑话,最终形成包含160,000余条样本的语料库。基于此,我们提出了首个专为中文幽默任务设计的大语言模型——中文趣模型(CFunModel),可完成相声回应选择、幽默识别、笑话生成等多类任务。实验表明,CFunModel在各项任务上均优于主流大模型。CFunSet数据集已公开于https://huggingface.co/datasets/ZhenghanYU/CFunSet,CFunModel模型资源可在https://huggingface.co/ZhenghanYU/CFunModel获取,演示视频见https://youtu.be/MOsISOJ66Ms。
原文摘要 · Abstract (English)
Humor plays a significant role in daily language communication. With the rapid development of large language models (LLMs), natural language processing has made significant strides in understanding and generating various genres of texts. However, most LLMs exhibit poor performance in generating and processing Chinese humor. In this study, we introduce a comprehensive Chinese humor-related dataset, the Chinese Fun Set (CFunSet). This dataset aggregates existing Chinese humor datasets and includes over 20,000 jokes collected from Tieba-JokeBar, a Chinese online platform known for joke sharing. The resulting corpus comprises more than 160,000 entries. Leveraging CFunSet, we developed the Chinese Fun Model (CFunModel), the first large language model designed to handle various Chinese humor-related tasks including Crosstalk Response Selection, Humor Recognition, Joke Generation, etc. Experimental results demonstrate that CFunModel outperforms popular large language models in these tasks. Our CFunSet is available at https://huggingface.co/datasets/ZhenghanYU/CFunSet and CFunModel is available at https://huggingface.co/ZhenghanYU/CFunModel. A demostration video of our work is available at https://youtu.be/MOsISOJ66Ms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。