用极少信息实现高效精准的假新闻检测,突破传统方法依赖全文的局限。
Is Less Really More? Fake News Detection with Limited Information
- 通过信息论量化信息量,系统选取最有效的少量信息进行检测。
- 仅用部分信息即达到顶尖模型的准确率,训练数据量大幅减少。
- 适合信息不全或资源受限场景,提升检测鲁棒性与实用性。
在线假新闻和误导性信息对民主、正义、公众信任及弱势群体构成严重威胁,推动了假新闻检测与干预的迫切需求。当前多数检测方法依赖文章全文的文本分析,存在计算效率低、需大量训练数据、跨数据集泛化能力差等问题。这是因为假新闻数据集在信息量和类型上差异显著:有的包含长段落、图片和元数据,有的仅由几句话组成。若能仅使用有限信息进行检测,可能使方法更稳健、更具适应性。为此,本文提出SLIM(Systematically-selected Limited Information)框架,引入信息论度量来量化信息量,通过系统选择有限信息实现与使用全文相当的检测性能。结合多种有限信息类型后,SLIM性能进一步提升,且所需训练信息量显著低于基于语言模型的主流方法。
原文摘要 · Abstract (English)
The threat that online fake news and misinformation pose to democracy, justice, public confidence, and especially to vulnerable populations, has led to a sharp increase in the need for fake news detection and intervention. Whether multi-modal or pure text-based, most fake news detection methods depend on textual analysis of entire articles. However, these fake news detection methods come with certain limitations. For instance, fake news detection methods that rely on full text can be computationally inefficient, demand large amounts of training data to achieve competitive accuracy, and may lack robustness across different datasets. This is because fake news datasets have strong variations in terms of the level and types of information they provide; where some can include large paragraphs of text with images and metadata, others can be a few short sentences. Perhaps if one could only use minimal information to detect fake news, fake news detection methods could become more robust and resilient to the lack of information. We aim to overcome these limitations by detecting fake news using systematically selected, limited information that is both effective and capable of delivering robust, promising performance. We propose a framework called SLIM Systematically-selected Limited Information) for fake news detection. In SLIM, we quantify the amount of information by introducing information-theoretic measures. SLIM leverages limited information to achieve performance in fake news detection comparable to that of state-of-the-art obtained using the full text. Furthermore, by combining various types of limited information, SLIM can perform even better while significantly reducing the quantity of information required for training compared to state-of-the-art language model-based fake news detection techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。