arXiv:2605.12510cs.SIcs.CL2026-05AAAI被引 1

构建首个专家标注的巴西疫苗谣言数据集,助力加密聊天中的健康假信息检测。

WhatsApp Vaccine Discourse (WhaVax): An Expert-Annotated Dataset and Benchmark for Health Misinformation Detection

论文配图:WhatsApp Vaccine Discourse (WhaVax): An Expert-Annotated Dataset and Benchmark for Health Misinformation Detection
图 1 · 摘自论文原文
  • 通过关键词抓取+语义去重+医学专家多阶段标注,确保数据高质量。
  • 发现私聊中假信息具有独特语言、时间与群组传播模式,含大量模糊案例。
  • 适合研究假信息传播、社交媒体安全与医疗沟通的学者及政策制定者。

我们提出WhaVax,一个从多个巴西公共群组收集的、覆盖疫情多年期的疫苗相关WhatsApp消息专家标注数据集。该数据集通过严谨的采集流程构建:基于关键词的数据收集、语义去重以消除近似重复内容,并由医疗专家执行多阶段标注,形成高可靠性的金标准语料库,具备显著的标注者间一致性。此外,我们对WhatsApp假信息进行了详细表征,揭示其在语言、结构、词汇、时间及群组层面的独特模式,并发现大量介于真假之间的模糊案例,反映出私密通信中健康话语的复杂性。我们在真实数据稀缺条件下评估了经典模型、微调的小型语言模型以及零样本或少样本的大语言模型,结果表明强嵌入与大模型方法表现优异,但领域适配性和数据可用性仍是关键影响因素。本研究为加密通信环境下的假信息研究与计算建模提供了罕见且高质量的资源。

原文摘要 · Abstract (English)

We introduce WhaVax, a new expert-annotated dataset of vaccine-related WhatsApp messages collected from large Brazilian public groups spanning multiple pandemic years. The dataset was constructed through a rigorous, carefully designed pipeline that integrates keyword-based data collection, semantic deduplication to remove near-duplicate content, and a multi-stage annotation protocol conducted by medical specialists. This process produced a high-quality gold-standard corpus, characterized by substantial inter-annotator agreement and strong reliability for downstream analysis. Additionally, we provide a detailed characterization of WhatsApp misinformation, revealing distinctive linguistic, structural, lexical, temporal, and group-level patterns, as well as a meaningful layer of ambiguous cases that reflect the complexity of health discourse in private messaging. We also benchmark classical models, fine-tuned Small Language Models, and zero- or few-shot Large Language Models under realistic data-scarcity constraints, demonstrating that strong embeddings and LLM approaches perform competitively, while domain alignment and data availability remain critical factors. This study provides a rare, high-quality resource to support misinformation research and computational modeling in encrypted communication environments.

假信息检测社交媒体医疗传播数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。