为孟加拉语构建大规模真实新闻数据集,提升低资源语言假新闻检测能力
From Scarcity to Capability: Empowering Fake News Detection in Low-Resource Languages with LLMs
- 构建包含47,000条真实与13,000条虚假新闻的孟加拉语数据集
- 基于大模型的检测系统在测试集上达F1-89%的准确率
- 开源数据集与模型,助力低资源语言假新闻研究
假新闻的快速传播构成全球性挑战,尤其在孟加拉语等低资源语言中,缺乏足够数据集与检测工具。尽管人工核验准确,但成本高、速度慢。为此,我们推出BanFakeNews-2.0,一个强化的孟加拉语假新闻数据集,新增11,700条经可信来源验证的虚假新闻,形成47,000条真实与13,000条虚假新闻的均衡数据集,覆盖13个类别。同时构建独立测试集,含460条虚假和540条真实新闻,用于严格评估。通过从可信来源收集并人工验证虚假新闻,保留语言丰富性。开发基于Transformer架构的基准系统,包括微调BERT变体(F1-87%)和量化低秩近似的大型语言模型(F1-89%),显著优于传统方法。我们公开发布数据集与模型于Github,推动低资源语言假新闻检测研究。
原文摘要 · Abstract (English)
The rapid spread of fake news presents a significant global challenge, particularly in low-resource languages like Bangla, which lack adequate datasets and detection tools. Although manual fact-checking is accurate, it is expensive and slow to prevent the dissemination of fake news. Addressing this gap, we introduce BanFakeNews-2.0, a robust dataset to enhance Bangla fake news detection. This version includes 11,700 additional, meticulously curated fake news articles validated from credible sources, creating a proportional dataset of 47,000 authentic and 13,000 fake news items across 13 categories. In addition, we created a manually curated independent test set of 460 fake and 540 authentic news items for rigorous evaluation. We invest efforts in collecting fake news from credible sources and manually verified while preserving the linguistic richness. We develop a benchmark system utilizing transformer-based architectures, including fine-tuned Bidirectional Encoder Representations from Transformers variants (F1-87\%) and Large Language Models with Quantized Low-Rank Approximation (F1-89\%), that significantly outperforms traditional methods. BanFakeNews-2.0 offers a valuable resource to advance research and application in fake news detection for low-resourced languages. We publicly release our dataset and model on Github to foster research in this direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。