针对乌尔都语假新闻检测,通过领域适配提升模型效果。
Fake News Classification in Urdu: A Domain Adaptation Approach for a Low-Resource Language
- 先用乌尔都语新闻数据做领域自适应预训练,再微调分类器。
- 领域适配后的XLM-R在四个数据集上均优于原始模型。
- 适合低资源语言假新闻检测研究者参考。
社交媒体上的虚假信息是一个全球性问题,但像乌尔都语这样的低资源语言在此领域关注较少。直接使用多语言预训练模型并微调虽可行,但对领域专有词汇处理不佳。为此,我们提出在微调前进行领域适配,采用分阶段训练策略提升模型泛化能力。评估了XLM-RoBERTa和mBERT两个主流多语言模型,利用公开的乌尔都语新闻语料进行领域自适应预训练。在四个公开的乌尔都语假新闻数据集上的实验表明,经过领域适配的XLM-R表现持续优于原始版本,而领域适配后的mBERT结果则不一致。
原文摘要 · Abstract (English)
Misinformation on social media is a widely acknowledged issue, and researchers worldwide are actively engaged in its detection. However, low-resource languages such as Urdu have received limited attention in this domain. An obvious approach is to utilize a multilingual pretrained language model and fine-tune it for a downstream classification task, such as misinformation detection. However, these models struggle with domain-specific terms, leading to suboptimal performance. To address this, we investigate the effectiveness of domain adaptation before fine-tuning for fake news classification in Urdu, employing a staged training approach to optimize model generalization. We evaluate two widely used multilingual models, XLM-RoBERTa and mBERT, and apply domain-adaptive pretraining using a publicly available Urdu news corpus. Experiments on four publicly available Urdu fake news datasets show that domain-adapted XLM-R consistently outperforms its vanilla counterpart, while domain-adapted mBERT exhibits mixed results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。