arXiv:2410.12061cs.SIcs.AI2024-10被引 4

结合知识库与社交网络,提升社交媒体假信息检测精度

CrediRAG: Network-Augmented Credibility-Based Retrieval for Misinformation Detection in Reddit

  • 融合政治知识库与用户互动网络,动态评估新闻可信度
  • 在超20万篇Reddit帖子上实现F1分数提升11%
  • 适合研究虚假信息传播、社交网络分析的学者与工程师

假新闻威胁民主并加剧社会分化,准确检测网络虚假信息是应对这一问题的基础。我们提出CrediRAG,首个将语言模型与丰富外部政治知识库结合,并引入密集社交网络以大规模检测社交媒体假新闻的模型。CrediRAG首先通过新闻检索器根据帖子标题内容匹配相似新闻文章的来源可信度,初步赋值伪信息分数;随后,基于共享评论者构建加权的帖子-帖子网络,权重由共享评论者的平均立场决定,进一步优化初始判断。在超过20万条真实Reddit数据上的实验表明,该方法相较现有最优方法实现F1分数11%的提升,验证了其在准确性与可扩展性上的优越性。本方法为应对社交媒体中假新闻传播提供了更精准、高效的新方案。

原文摘要 · Abstract (English)

Fake news threatens democracy and exacerbates the polarization and divisions in society; therefore, accurately detecting online misinformation is the foundation of addressing this issue. We present CrediRAG, the first fake news detection model that combines language models with access to a rich external political knowledge base with a dense social network to detect fake news across social media at scale. CrediRAG uses a news retriever to initially assign a misinformation score to each post based on the source credibility of similar news articles to the post title content. CrediRAG then improves the initial retrieval estimations through a novel weighted post-to-post network connected based on shared commenters and weighted by the average stance of all shared commenters across every pair of posts. We achieve 11% increase in the F1-score in detecting misinformative posts over state-of-the-art methods. Extensive experiments conducted on curated real-world Reddit data of over 200,000 posts demonstrate the superior performance of CrediRAG on existing baselines. Thus, our approach offers a more accurate and scalable solution to combat the spread of fake news across social media platforms.

假信息检测社交网络知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。