arXiv:2503.16031cs.CL2025-03被引 4

构建多语言讽刺假消息数据集,揭示幽默如何掩盖虚假信息

Deceptive Humor: A Synthetic Multilingual Benchmark Dataset for Bridging Fabricated Claims with Humorous Content

  • 用ChatGPT-4o生成带幽默的假言论,标注讽刺等级与五类幽默类型
  • 涵盖英、印地、泰米尔等六种语言及代码混杂形式,覆盖广泛语境
  • 为识别幽默化谣言提供基准,适合研究虚假信息传播与检测的学者

在在线话语演进中,虚假信息越来越多地采用幽默语气以逃避检测并扩大传播。本文提出‘欺骗性幽默’这一新研究方向,强调当错误叙事披上幽默外衣时,更难被识别且更易扩散。为此,我们构建了欺骗性幽默数据集(DHD),包含使用ChatGPT-4o生成的幽默化评论,每条标注了讽刺等级(1为微妙讽刺,3为明显讽刺)并归类为五种幽默类型:黑色幽默、反语、社会评论、文字游戏和荒诞。数据集覆盖英语、泰卢固语、印地语、卡纳达语、泰米尔语及其代码混杂形式,支持多语言分析。DHD为理解幽默如何成为虚假信息传播载体提供了结构化基础,显著增强其传播力与影响力。同时建立了强基线模型,推动该新兴领域研究与模型发展。

原文摘要 · Abstract (English)

In the evolving landscape of online discourse, misinformation increasingly adopts humorous tones to evade detection and gain traction. This work introduces Deceptive Humor as a novel research direction, emphasizing how false narratives, when coated in humor, can become more difficult to detect and more likely to spread. To support research in this space, we present the Deceptive Humor Dataset (DHD) a collection of humor-infused comments derived from fabricated claims using the ChatGPT-4o model. Each entry is labeled with a Satire Level (from 1 for subtle satire to 3 for overt satire) and categorized into five humor types: Dark Humor, Irony, Social Commentary, Wordplay, and Absurdity. The dataset spans English, Telugu, Hindi, Kannada, Tamil, and their code-mixed forms, making it a valuable resource for multilingual analysis. DHD offers a structured foundation for understanding how humor can serve as a vehicle for the propagation of misinformation, subtly enhancing its reach and impact. Strong baselines are established to encourage further research and model development in this emerging area.

虚假信息幽默识别多语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。