arXiv:2605.01292cs.CL2026-05中稿 · 15th ACM ICSCA, 20…被引 2

用大模型生成假新闻数据,提升孟加拉语虚假信息检测效果

Addressing Data Scarcity in Bangla Fake News Detection: An LLM-Based Dataset Augmentation Approach

论文配图:Addressing Data Scarcity in Bangla Fake News Detection: An LLM-Based Dataset Augmentation Approach
图 1 · 摘自论文原文
  • 用微调过的Gemma 3 27B模型生成合成孟加拉语新闻,结合语义过滤与采样控制
  • 仅对少数类高倍率随机增广,使假新闻F1值从0.85提升至0.88
  • 适合关注低资源语言、伪造内容检测与大模型数据增强的研究者

数字媒体中虚假信息的蔓延凸显了可靠假新闻检测系统的需求,但孟加拉语等低资源语言因数据量小且不平衡而进展受限。本研究探究基于大语言模型(LLM)的数据增强是否能有效缓解这一问题并提升孟加拉语假新闻分类性能。现有数据集虽有价值,但严重失衡,制约模型表现,而针对孟加拉语的LLM数据增强尚未被充分探索。为此,我们提出一个系统性增强框架,利用指令微调的Gemma 3 27B IT模型生成合成孟加拉语新闻文章,并通过语义过滤和可控子采样确保标签一致性和多样性。我们比较了零样本与少样本提示策略,评估多种增强率,并分析随机与基于相似性的选择策略。实验表明,仅对少数类进行高倍率增强并采用随机子采样可取得最佳效果,使假新闻F1得分从0.85提升至0.88。为支持可复现性与进一步研究,我们公开发布4,545个合成孟加拉语假新闻样本及完整实现代码。结果表明,精心设计的LLM驱动增强可显著提升低资源场景下的假新闻检测能力,为多语言虚假信息研究提供实用基础。

原文摘要 · Abstract (English)

The growing spread of misinformation in digital media highlights the need for reliable fake news detection systems, yet progress in under-resourced languages such as Bangla is limited by small and imbalanced datasets. This study investigates whether Large Language Model (LLM) based augmentation can effectively address this limitation and improve Bangla fake news classification. Existing datasets remain valuable but highly imbalanced, limiting model performance, and LLM based augmentation for Bangla has been scarcely explored. To fill this gap, we propose a systematic augmentation framework that generates synthetic Bangla news articles using the instruction tuned Gemma 3 27B IT model, supported by semantic filtering and controlled subsampling to preserve label consistency and diversity. We compare zero shot and few shot prompting, evaluate multiple augmentation rates, and examine random versus similarity-based selection strategies. Our experiments show that augmenting only the minority class with a high augmentation rate and random subsampling yields the strongest gains, raising the Fake News F1 score from 0.85 to 0.88. To support reproducibility and further research in this low-resource domain, we publicly release 4,545 synthetically generated Bangla fake news samples along with our full implementation. These findings demonstrate that well-designed LLM-driven augmentation can significantly improve fake news detection in low resource settings and provide a practical foundation for advancing multilingual misinformation research.

假新闻检测大模型数据增强低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。