arXiv:2608.02399cs.SIcs.LG2026-08

利用社交媒体分享网络结构,提升虚假新闻源识别准确率

Network Information Enhances Unreliable News Domain Detection

论文配图:Network Information Enhances Unreliable News Domain Detection
图 1 · 摘自论文原文
  • 构建域名共分享网络,发现低可信与高可信域名各自聚集
  • 图神经网络相比传统模型提升13%-14%准确率,最高达63%
  • 无需内容分析也能有效识别,适合内容不可靠场景

基于内容的虚假新闻检测日益困难,因低可信度来源模仿正规新闻,且生成式AI使伪造内容更难识别。本文提出从域级别出发,关注新闻源可靠性而非单篇文章。通过分析Telegram聊天中的网址分享模式,构建经统计验证的域名共分享网络,发现可靠性存在同质性聚集:低可信域名与高可信域名分别聚类。利用该结构,对比图神经网络(GNN)与无网络感知基线模型,使用多语言文本嵌入(内容相关特征)和传播动态(内容无关特征)。在相同特征下,GNN始终优于多层感知机,GraphSAGE表现最佳(含内容特征时准确率0.63,不含时0.53),相对基线提升13%-14%。结果表明,网络拓扑结构能系统性提升域级可靠性评估,即使内容分析不可行也有效。

原文摘要 · Abstract (English)

Content-based detection of unreliable news is increasingly difficult, as low-reliability sources mimic credible journalism and generative AI makes fabricated content harder to flag. We ask whether network structure can improve news reliability classification, taking a domain-level approach that shifts the focus from individual articles to source reliability. From URL-sharing patterns in Telegram chats, we build a statistically validated domain co-sharing network and find assortative mixing by reliability: low-reliability domains group together, as do reliable ones. Exploiting this structure, we compare Graph Neural Networks against network-unaware baselines using both content-aware features (multilingual text embeddings) and content-agnostic features (spreading dynamics). GNNs consistently outperform Multi-Layer Perceptrons on identical features, with GraphSAGE best in both settings (accuracy 0.63 with content, 0.53 without), a 13-14% relative gain over the network-unaware baseline. Network topology thus systematically improves domain reliability assessment, and remains effective even when content analysis is infeasible.

虚假新闻图神经网络社交网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。