arXiv:2605.06435cs.CLcs.AI2026-05被引 6

用文本和语言特征检测新冠假新闻,传统机器学习有效

COVID-19 Infodemic. Understanding content features in detecting fake news using a machine learning approach

  • 提取词组、词性分布等文本特征进行假新闻识别
  • 随机森林模型准确率最高,达87.3%
  • 文本与语言特征单独使用效果好,合并提升不明显

尽管实证研究表明文本与语言特征有助于区分真假新闻,但其在假新闻检测中的应用仍研究不足。本研究针对新冠疫情期间的数据集,考察词组、词性分布等内容特征对假新闻检测的影响。实验采用决策树、K近邻、逻辑回归、支持向量机和随机森林五种传统机器学习模型。结果显示,随机森林表现最佳,准确率达87.3%,紧随其后的是支持向量机。无论是文本特征还是语言特征单独使用,均能提升检测性能;但二者合并后效果提升不显著。此外,词组与词性标签的使用效果存在差异。研究证实,传统机器学习可有效利用文本与语言特征实现假新闻检测。

原文摘要 · Abstract (English)

The use of content features, particularly textual and linguistic for fake news detection is under-researched, despite empirical evidence showing the features could contribute to differentiating real and fake news. To this end, this study investigates a selection of content features such as word bigrams, part of speech distribution etc. to improve fake news detection. We performed a series of experiments on a new dataset gathered during the COVID-19 pandemic and using Decision Tree, K-Nearest Neighbor, Logistic Regression, Support Vector Machine and Random Forest. Random Forest yielded the best results, followed closely by Support Vector Machine, across all setups. In general, both the textual and linguistic features were found to improve fake news detection when used separately, however, combining them into a single model did not improve the detection significantly. Differences were also noted between the use of bigrams and part of speech tags. The study shows that textual and linguistic features can be used successfully in detecting fake news using the traditional machine learning approach as opposed to deep learning.

假新闻检测文本特征机器学习新冠信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。