arXiv:2605.17481cs.CL2026-05中稿 · and presented in t…

用混合特征提升孟加拉语假新闻识别准确率

Hybrid Feature Combinations with CNN for Bangla Fake News Classification

论文配图:Hybrid Feature Combinations with CNN for Bangla Fake News Classification
图 1 · 摘自论文原文
  • 融合语义、统计与字符级特征提升检测效果
  • 多特征组合使召回率和F1分数显著提高
  • 适合关注低资源语言假新闻检测的研究者

如今,孟加拉国人越来越依赖互联网和社交媒体获取日常新闻,而非传统报纸。然而,虚假孟加拉语新闻在这些平台上的传播对真实媒体的可信度构成威胁。尽管已有研究致力于检测孟加拉语假新闻,但仍有改进空间。本研究基于BanFakeNews-2.0数据集,探索语义、统计及字符级特征或其组合在卷积神经网络(CNN)模型中的有效性。结果表明,多特征组合显著优于单一特征,能有效提升召回率与F1分数。相关代码已公开于GitHub:https://github.com/gulzar09/Bn_FNews_H.Feature。

原文摘要 · Abstract (English)

Nowadays, people in Bangladesh frequently rely on the internet and social media for daily news instead of traditional newspapers. However, the spread of false Bangla news through these platforms poses risks and challenges to the credibility of authentic media. Although several studies have been conducted on detecting Bangla fake news, there is still significant room for improvement in this area. To assist people, this research explores the effectiveness of feature selection approaches in identifying appropriate features, such as semantic, statistical, and character-level features, or their combinations, on the BanFakeNews-2.0 dataset for detecting Bangla fake news using a CNN model. In this paper, key findings reveal that combining multiple features significantly improves recall and F1-scores compared to using individual features alone. The code for this research can be availed here, https://github.com/gulzar09/Bn\_FNews\_H.Feature.

假新闻检测深度学习孟加拉语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。