用BERT提升孟加拉语极端偏见新闻识别准确率
Bangla BERT for Hyperpartisan News Detection: A Semi-Supervised and Explainable AI Approach
- 微调孟加拉语BERT模型,结合半监督学习提升分类性能
- 达到95.65%准确率,显著优于传统机器学习方法
- 使用LIME提供决策解释,增强模型可信度
在当前数字环境中,虚假信息传播迅速,影响公众认知并加剧社会分裂。由于缺乏针对低资源语言孟加拉语的先进自然语言处理方法,识别极端偏见新闻尤为困难。若无有效检测手段,偏见内容将难以遏制,威胁理性讨论。为此,本研究对孟加拉语BERT进行微调,这是一种基于Transformer的先进模型,旨在提升极端偏见新闻分类的准确性。实验对比了传统机器学习模型,并引入半监督学习进一步优化预测效果。同时,采用LIME技术提供模型决策过程的透明解释,增强结果可信度。测试结果显示,该方法准确率达95.65%,显著优于传统方法。研究证明,即使在资源有限环境下,变压器模型仍具强大实用性,为后续改进开辟路径。
原文摘要 · Abstract (English)
In the current digital landscape, misinformation circulates rapidly, shaping public perception and causing societal divisions. It is difficult to identify hyperpartisan news in Bangla since there aren't many sophisticated natural language processing methods available for this low-resource language. Without effective detection methods, biased content can spread unchecked, posing serious risks to informed discourse. To address this gap, our research fine-tunes Bangla BERT. This is a state-of-the-art transformer-based model, designed to enhance classification accuracy for hyperpartisan news. We evaluate its performance against traditional machine learning models and implement semi-supervised learning to enhance predictions further. Not only that, we use LIME to provide transparent explanations of the model's decision-making process, which helps to build trust in its outcomes. With a remarkable accuracy score of 95.65%, Bangla BERT outperforms conventional approaches, according to our trial data. The findings of this study demonstrate the usefulness of transformer models even in environments with limited resources, which opens the door to further improvements in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。