构建首个面向大模型时代的假新闻检测图文图数据集
TAGFN: A Text-Attributed Graph Dataset for Fake News Detection in the Age of LLMs
- 构建真实世界图文图数据集,支持假新闻等异常检测
- 涵盖大规模真实数据,可评估传统与大模型方法
- 适合研究假新闻检测、可信AI与大模型微调的学者
大型语言模型(LLMs)在图文图任务中已取得突破,但在图异常检测(如假新闻检测)中的应用仍严重不足。主要瓶颈在于缺乏大规模、真实且标注完善的基准数据集。为此,我们提出TAGFN——一个大规模、真实世界的文本属性图数据集,专为异常检测(尤其是假新闻检测)设计。该数据集支持对传统方法与基于大模型的图异常检测方法进行严格评估,并可用于通过微调提升大模型在虚假信息识别方面的能力。我们相信,TAGFN将推动鲁棒图异常检测与可信人工智能的发展。数据集已公开于https://huggingface.co/datasets/kayzliu/TAGFN,代码见https://github.com/kayzliu/tagfn。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have recently revolutionized machine learning on text-attributed graphs, but the application of LLMs to graph outlier detection, particularly in the context of fake news detection, remains significantly underexplored. One of the key challenges is the scarcity of large-scale, realistic, and well-annotated datasets that can serve as reliable benchmarks for outlier detection. To bridge this gap, we introduce TAGFN, a large-scale, real-world text-attributed graph dataset for outlier detection, specifically fake news detection. TAGFN enables rigorous evaluation of both traditional and LLM-based graph outlier detection methods. Furthermore, it facilitates the development of misinformation detection capabilities in LLMs through fine-tuning. We anticipate that TAGFN will be a valuable resource for the community, fostering progress in robust graph-based outlier detection and trustworthy AI. The dataset is publicly available at https://huggingface.co/datasets/kayzliu/TAGFN and our code is available at https://github.com/kayzliu/tagfn.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。