arXiv:2509.23340cs.SIcs.DC2025-09KDD

构建大规模网页可信度数据集,融合链接、文本与时间信息提升假信息识别能力。

CrediBench: Building Web-Scale Network Datasets for Information Integrity

  • 基于8个月网页图数据,融合链接、文本与时间多模态信号
  • 涵盖超4000万节点和10亿级超链接,支持回归与分类任务
  • 模型性能显著提升:分类准确率从56%增至85%

自动评估网络来源可信度是应对当前信息生态挑战的重要工具。现有方法或依赖稀缺人力标注,或仅关注单个声明层面。虚假信息常通过相互关联的网页域传播,其关系随时间演变。仅关注声明会忽略网页拓扑结构与时间动态等关键信号。现有数据集未能涵盖网络拓扑、时间性和文本内容这三大核心维度。为此,我们提出CrediBench,一个包含8个月网页图数据的数据集,重点分析2024年美国联邦选举期间的三个月,该时期线上虚假信息传播尤为频繁。每月快照包含超过4000万节点、抓取的网页内容及超10亿条超链接边。该数据集支持可信度预测的回归(连续可信分数)与二分类任务(可信/不可信)。我们构建了包含662,575个网页域的新型二元标签集,覆盖虚假信息、众包、恶意软件和钓鱼四大类别。实验证明,图、文本与时间三类模态均显著贡献于最优性能。特别是,基于CrediBench训练的多模态回归模型将平均误差从0.162降至0.107;多模态分类器准确率从56%提升至85%。CrediBench已开源至Huggingface,供后续研究使用。

原文摘要 · Abstract (English)

Automatically assessing the credibility of online sources presents an invaluable tool for navigating today's information ecosystem. However, existing approaches either depend on scarce and costly human annotations, or focus exclusively on assessments at the level of individual claims. Misinformation often spreads via interlinked web domains, whose connections evolve over time. Focusing on claims alone ignores these structural and temporal credibility signals evident in the changing web topology. Existing datasets fail to capture these central modalities in web domain credibility prediction: namely, internet topology, temporality and text (webpage) content. To address this gap, we present CrediBench, a dataset containing eights months of web graph data; of which we analyze the three months surrounding the 2024 U.S. federal elections, a time of heightened misinformation propagation online. Each monthly snapshot contains over 40 million nodes, their scraped webpage content, and over 1 billion hyperlink edges. CrediBench supports credibility prediction as both a regression (continuous credibility score) and a binary classification task (credible or not). For classification, we curate a novel binary label set containing 662,575 web domains labelled for boolean credibility, spanning four areas (misinformation, crowd-sourced, malware and phishing). Our empirical experiments support that all task modalities-graph, text and time-contribute significantly to achieving the best performance. In particular, our multi-modal regression model trained on CrediBench outperforms other configurations and existing baselines, decreasing Mean Average Error from 0.162 to 0.107 on the regression task, while the multi-modal classifier improves accuracy from 56% to 85% on the classification one. CrediBench, our proposed web-scale multi-modal dataset, is available on Huggingface for future research.

可信度评估多模态网络分析数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。