首个公开的乌尔都语假新闻检测数据集,助力低资源语言信息甄别。
Unified Large Language Models for Misinformation Detection in Low-Resource Linguistic Settings
- 构建首个公开可用的乌尔都语假新闻检测基准数据集。
- 统一模型在准确率与F1值上优于XLNet、mBERT等主流大模型。
- 数据集经专家验证,适用于低资源语言虚假信息研究。
社交媒体快速扩张加剧了虚假内容传播,假新闻检测成为关键研究方向。尽管现有事实核查多集中于英语,但区域性语言如乌尔都语仍缺乏有效资源与策略。先进假新闻检测依赖大规模高质量标注数据,而乌尔都语等低资源语言因语料稀缺且缺乏验证词汇资源,面临严峻挑战。现有乌尔都语数据集多为领域特定、未公开,且主要依赖未经验证的英译乌尔都语,可靠性存疑。本研究强调需建立可靠、专家验证、跨领域的乌尔都语增强型假新闻检测数据集。本文提出首个公开可验证的大型乌尔都语假新闻检测基准数据集,并使用XLNet、mBERT、XLM-RoBERTa、RoBERTa、DistilBERT和DeBERTa等主流预训练大语言模型进行评估。此外,我们提出一种统一的大模型架构,在不同嵌入与特征提取技术下表现更优,模型性能通过准确率、F1分数、精确率、召回率及人工审核样本结果综合评估。
原文摘要 · Abstract (English)
The rapid expansion of social media platforms has significantly increased the dissemination of forged content and misinformation, making the detection of fake news a critical area of research. Although fact-checking efforts predominantly focus on English-language news, there is a noticeable gap in resources and strategies to detect news in regional languages, such as Urdu. Advanced Fake News Detection (FND) techniques rely heavily on large, accurately labeled datasets. However, FND in under-resourced languages like Urdu faces substantial challenges due to the scarcity of extensive corpora and the lack of validated lexical resources. Current Urdu fake news datasets are often domain-specific and inaccessible to the public. They also lack human verification, relying mainly on unverified English-to-Urdu translations, which compromises their reliability in practical applications. This study highlights the necessity of developing reliable, expert-verified, and domain-independent Urdu-enhanced FND datasets to improve fake news detection in Urdu and other resource-constrained languages. This paper presents the first benchmark large FND dataset for Urdu news, which is publicly available for validation and deep analysis. We also evaluate this dataset using multiple state-of-the-art pre-trained large language models (LLMs), such as XLNet, mBERT, XLM-RoBERTa, RoBERTa, DistilBERT, and DeBERTa. Additionally, we propose a unified LLM model that outperforms the others with different embedding and feature extraction techniques. The performance of these models is compared based on accuracy, F1 score, precision, recall, and human judgment for vetting the sample results of news.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。