arXiv:2609.02895cs.CLcs.LG2026-09

构建印度公共事件假信息检测数据集,填补文化语境空白

BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events

论文配图:BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events
图 1 · 摘自论文原文
  • 融合事实核查平台、多媒体转录与大模型生成,构建多源数据集
  • 包含14,646条记录,支持事件感知的假信息二分类任务
  • 专为印度宗教、政治等大型活动设计,适合本地化内容安全研究

大型公共事件如宗教节庆、政治集会和文化聚会正面临虚假信息快速传播的风险,威胁公共安全与社会团结。尽管自动化假新闻检测已取得显著进展,现有基准难以捕捉印度语境下的社会文化特征与事件动态。本文提出BharatGather,一个针对印度大规模集会场景的二元假信息分类专用数据集。该语料库由14,646条记录构成,通过混合管道整合主流事实核查平台的系统性网络爬取、多媒体转录提取及大语言模型(LLM)辅助的合成数据增强,保障叙事多样性。该资源旨在推动具备文化敏感性的检测系统研发,并为高风险公共环境中模型性能评估提供严格基准。

原文摘要 · Abstract (English)

Large-scale public events, such as religious festivals, political rallies, and cultural gatherings, are increasingly vulnerable to the rapid dissemination of misinformation, posing substantial risks to public safety and social cohesion. While automated fake news detection has seen significant methodological progress, existing benchmarks frequently fail to capture the socio-cultural nuances and event-specific dynamics characteristic of the Indian context. This paper introduces BharatGather, a curated, multi-source dataset specifically engineered for binary misinformation classification within the ecosystem of Indian mass gatherings. The corpus comprises 14,646 records constructed through a hybrid pipeline involving systematic web scraping of prominent fact-checking platforms, multimedia transcript extraction, and Large Language Model (LLM)-mediated synthetic augmentation to ensure narrative diversity. By providing a resource tailored to the unique complexities of event-aware misinformation in India, this work facilitates the development of culturally informed detection systems and establishes a rigorous benchmark for evaluating their performance in high-stakes public environments.

假新闻检测数据集印度文化敏感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。