arXiv:2412.18618cs.CLcs.SI2024-12

通过主题一致性差异识别假新闻文本特征

Exploring Text Representations for Online Misinformation

  • 利用真伪新闻在主题演变上的差异提取文本特征
  • 主题特征在分类与聚类中均有效提升检测性能
  • 无需标注数据即可聚类,适合小样本场景

虚假信息(包括错误信息和误导性信息)持续对社会构成威胁,尤其在政治和医疗领域影响显著,并正扩展至更多领域。尽管其传播形式多样,但主要以文本形式在社交媒体和博客文章中扩散。因此,尽早检测文本类虚假信息至关重要。本文提出一种新方法,从新闻文章中提取用于检测虚假信息的文本特征,核心思路是利用真实新闻与虚假新闻在主题连贯性上的差异——随着故事推进,两类新闻所讨论的主题组合存在显著不同。实验表明,该主题特征在分类与聚类任务中均表现良好;其中聚类方法尤其有价值,因其可不依赖人工标注数据,缓解了标注成本高、耗时长的问题。本研究推动了对虚假信息本质的理解,为基于机器学习与自然语言处理的检测提供了新思路。

原文摘要 · Abstract (English)

Mis- and disinformation, commonly collectively called fake news, continue to menace society. Perhaps, the impact of this age-old problem is presently most plain in politics and healthcare. However, fake news is affecting an increasing number of domains. It takes many different forms and continues to shapeshift as technology advances. Though it arguably most widely spreads in textual form, e.g., through social media posts and blog articles. Thus, it is imperative to thwart the spread of textual misinformation, which necessitates its initial detection. This thesis contributes to the creation of representations that are useful for detecting misinformation. Firstly, it develops a novel method for extracting textual features from news articles for misinformation detection. These features harness the disparity between the thematic coherence of authentic and false news stories. In other words, the composition of themes discussed in both groups significantly differs as the story progresses. Secondly, it demonstrates the effectiveness of topic features for fake news detection, using classification and clustering. Clustering is particularly useful because it alleviates the need for a labelled dataset, which can be labour-intensive and time-consuming to amass. More generally, it contributes towards a better understanding of misinformation and ways of detecting it using Machine Learning and Natural Language Processing.

虚假信息文本检测主题建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。