用GRU模型识别孟加拉语假新闻,准确率达94%。
Breaking the Fake News Barrier: Deep Learning Approaches in Bangla Language
- 采用GRU网络结合分词、词形还原等预处理技术
- 在58,478条数据上实现94%准确率
- 首个大规模孟加拉语假新闻数据集与模型
数字平台的快速发展加剧了虚假信息的传播,尤其对讲孟加拉语的群体造成认知冲击。本文提出一种基于门控循环单元(GRU)的深度学习方法,用于识别孟加拉语假新闻。研究通过密集的数据预处理,包括词形还原、分词及过采样以应对数据不平衡问题,构建了一个包含58,478条文本的大型数据集。所提出的GRU模型在准确率、精确率、召回率和F1分数等指标上表现优异,准确率达到94%。该工作不仅提供了完整的数据处理与建模流程,还贡献了首个大规模孟加拉语假新闻数据集,其性能优于现有其他孟加拉语假新闻检测模型。
原文摘要 · Abstract (English)
The rapid development of digital stages has greatly compounded the dispersal of untrue data, dissolving certainty and judgment in society, especially among the Bengali-speaking community. Our ponder addresses this critical issue by presenting an interesting strategy that utilizes a profound learning innovation, particularly the Gated Repetitive Unit (GRU), to recognize fake news within the Bangla dialect. The strategy of our proposed work incorporates intensive information preprocessing, which includes lemmatization, tokenization, and tending to course awkward nature by oversampling. This comes about in a dataset containing 58,478 passages. We appreciate the creation of a demonstration based on GRU (Gated Repetitive Unit) that illustrates remarkable execution with a noteworthy precision rate of 94%. This ponder gives an intensive clarification of the methods included in planning the information, selecting the show, preparing it, and assessing its execution. The performance of the model is investigated by reliable metrics like precision, recall, F1 score, and accuracy. The commitment of the work incorporates making a huge fake news dataset in Bangla and a demonstration that has outperformed other Bangla fake news location models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。