通过挖掘常识冲突提升谣言检测鲁棒性,无需修改主模型
Robust Misinformation Detection by Visiting Potential Commonsense Conflict
- 用COMET提取文章常识三元组,对比真实与生成差异来构建冲突表达
- 在4个公开数据集和新数据集CoMis上均显著优于现有方法
- 可作为插件式增强模块,适用于各类谣言检测模型
互联网技术的发展导致虚假信息泛滥,对多个领域造成严重负面影响。为应对这一挑战,自动识别网络虚假信息的谣言检测(MD)成为研究热点。本文提出一种新型即插即用的增强方法——基于潜在常识冲突的谣言检测(MD-PCC)。受先前研究启发,假新闻更可能包含常识冲突。因此,我们利用COMET这一成熟常识推理工具,提取文章的常识三元组,并通过对比其与黄金标准三元组的差异,构建潜在常识冲突表达,作为每篇文章的增强信息。这些表达可直接用于扩充训练数据,任何现有MD模型均可在此基础上训练。此外,我们还构建了一个新的常识导向数据集CoMis,其中所有虚假新闻均由常识冲突引发。将MD-PCC集成到多种主流MD模型中,在4个公开基准数据集和CoMis上进行评估,结果表明该方法能持续超越现有基线。
原文摘要 · Abstract (English)
The development of Internet technology has led to an increased prevalence of misinformation, causing severe negative effects across diverse domains. To mitigate this challenge, Misinformation Detection (MD), aiming to detect online misinformation automatically, emerges as a rapidly growing research topic in the community. In this paper, we propose a novel plug-and-play augmentation method for the MD task, namely Misinformation Detection with Potential Commonsense Conflict (MD-PCC). We take inspiration from the prior studies indicating that fake articles are more likely to involve commonsense conflict. Accordingly, we construct commonsense expressions for articles, serving to express potential commonsense conflicts inferred by the difference between extracted commonsense triplet and golden ones inferred by the well-established commonsense reasoning tool COMET. These expressions are then specified for each article as augmentation. Any specific MD methods can be then trained on those commonsense-augmented articles. Besides, we also collect a novel commonsense-oriented dataset named CoMis, whose all fake articles are caused by commonsense conflict. We integrate MD-PCC with various existing MD backbones and compare them across both 4 public benchmark datasets and CoMis. Empirical results demonstrate that MD-PCC can consistently outperform the existing MD baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。