针对乌尔都语假新闻,区分人类与机器生成内容。
Detection of Human and Machine-Authored Fake News in Urdu
- 构建分层检测框架,同时识别真假新闻与生成主体。
- 在四个数据集上验证有效,提升低资源语言检测性能。
- 填补中文外语假新闻检测空白,适合多语言研究者。
社交媒体的兴起加剧了假新闻的传播,而像ChatGPT这样的大语言模型使生成高度可信、无错误的虚假信息变得更容易,使公众更难辨别真伪。依赖语言特征的传统假新闻检测方法也逐渐失效。此外,现有检测器主要关注二分类任务和英文文本,常忽略机器生成的真实与虚假新闻的区别,以及低资源语言中的检测问题。为此,本文将检测范式扩展至包含机器生成新闻,并聚焦于乌尔都语。提出一种分层检测策略以提升准确率与鲁棒性。实验表明该方法在四种不同设置下的数据集上均表现有效。
原文摘要 · Abstract (English)
The rise of social media has amplified the spread of fake news, now further complicated by large language models (LLMs) like ChatGPT, which ease the generation of highly convincing, error-free misinformation, making it increasingly challenging for the public to discern truth from falsehood. Traditional fake news detection methods relying on linguistic cues also becomes less effective. Moreover, current detectors primarily focus on binary classification and English texts, often overlooking the distinction between machine-generated true vs. fake news and the detection in low-resource languages. To this end, we updated detection schema to include machine-generated news with focus on the Urdu language. We further propose a hierarchical detection strategy to improve the accuracy and robustness. Experiments show its effectiveness across four datasets in various settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。