通过风格与经济激励特征提升假新闻检测模型的跨数据集泛化能力
An exploration of features to improve the generalisability of fake news detection models
- 提取文本风格和社交媒体变现特征,降低对标注偏差的敏感度
- 在NELA和Facebook URLs数据集上验证,新特征显著优于传统模型
- 适合关注真实场景下假新闻检测鲁棒性的研究者与实践者
假新闻在全球范围内影响选举与信息传播,检测至关重要。现有NLP与监督学习方法在交叉验证中表现良好,但在不同数据集间泛化能力差,即使同领域也如此。这源于粗粒度标注数据——文章按发布源标记,引入偏见,导致基于词元的模型(如TF-IDF、BERT)敏感于此类偏差。尽管大语言模型(LLMs)有潜力,但其在该领域的应用仍有限。本研究证明,即使在粗粒度标注数据中,仍可提取有意义特征以增强实际应用的鲁棒性。重点探索了词汇、句法与语义等风格特征,因其对数据集偏见敏感度较低;并引入新颖的社会-变现特征,捕捉假新闻背后的经济动机,如广告、外链与社交元素。模型在粗粒度标注的NELA 2020-21数据集上训练,使用人工标注的Facebook URLs数据集(假新闻检测领域金标准)进行评估。结果表明,基于词元的模型在有偏数据上表现受限,且缺乏对LLaMA等大模型在此任务上的有效证据。分析显示,风格与社会-变现特征相比词元方法及大模型更具泛化能力。统计与置换特征重要性分析进一步揭示其提升性能与缓解数据偏见的潜力,为改进假新闻检测提供了可行路径。
原文摘要 · Abstract (English)
Fake news poses global risks by influencing elections and spreading misinformation, making detection critical. Existing NLP and supervised Machine Learning methods perform well under cross-validation but struggle to generalise across datasets, even within the same domain. This issue stems from coarsely labelled training data, where articles are labelled based on their publisher, introducing biases that token-based models like TF-IDF and BERT are sensitive to. While Large Language Models (LLMs) offer promise, their application in fake news detection remains limited. This study demonstrates that meaningful features can still be extracted from coarsely labelled data to improve real-world robustness. Stylistic features-lexical, syntactic, and semantic-are explored due to their reduced sensitivity to dataset biases. Additionally, novel social-monetisation features are introduced, capturing economic incentives behind fake news, such as advertisements, external links, and social media elements. The study trains on the coarsely labelled NELA 2020-21 dataset and evaluates using the manually labelled Facebook URLs dataset, a gold standard for generalisability. Results highlight the limitations of token-based models trained on biased data and contribute to the scarce evidence on LLMs like LLaMa in this field. Findings indicate that stylistic and social-monetisation features offer more generalisable predictions than token-based methods and LLMs. Statistical and permutation feature importance analyses further reveal their potential to enhance performance and mitigate dataset biases, providing a path forward for improving fake news detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。