arXiv:2412.09203cs.CL2024-12被引 6

构建去毒幽默数据集,提升生成幽默的可靠性和安全性。

CleanComedy: Creating Friendly Humor through Generative Techniques

  • 构建中英文双语去毒笑话数据集,过滤毒性内容。
  • 对比人类与生成笑话的幽默度和毒性水平,验证数据质量。
  • 适合研究幽默生成、内容安全与多语言自然语言处理者。

由于资源有限且现有数据集质量不佳,幽默生成在自然语言处理中仍具挑战性。现有幽默语料常含毒性与重复内容,制约模型训练效果。本文提出 CleanComedy,一个经过部分标注、经毒性过滤的英俄双语笑话语料库,数据来自多个来源。通过幽默与毒性水平的调查,评估不同笑话组的数据过滤效果。同时,比较人类创作与多种生成模型(包括基于 CleanComedy 训练的基线模型)生成的笑话,探索计算机幽默生成的进展。

原文摘要 · Abstract (English)

Humor generation is a challenging task in natural language processing due to limited resources and the quality of existing datasets. Available humor language resources often suffer from toxicity and duplication, limiting their effectiveness for training robust models. This paper proposes CleanComedy, a specialized, partially annotated toxicity-filtered corpus of English and Russian jokes collected from various sources. We study the effectiveness of our data filtering approach through a survey on humor and toxicity levels in various joke groups. In addition, we study advances in computer humor generation by comparing jokes written by humans with various groups of generative jokes, including our baseline models trained on the CleanComedy datasets.

幽默生成数据过滤多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。