通过自适应对抗训练提升假新闻检测对恶意评论的鲁棒性
Group-Adaptive Adversarial Learning for Robust Fake News Detection Against Malicious Comments
- 按认知心理学分类恶意评论并用大模型生成针对性扰动
- 动态调整攻击比例,聚焦模型薄弱区域进行优化
- 在三个数据集上F1提升最高达17.9%,显著增强抗攻击能力
在线假新闻严重扭曲公众判断并侵蚀社交平台信任。尽管现有检测器在基准数据集上表现良好,但仍易受专门设计的恶意评论诱导误判。当前检测器难以泛化到多样且新颖的评论攻击模式。为此,我们提出AdComment框架,一种用于增强对多样化恶意评论鲁棒性的自适应对抗训练方法。基于认知心理学,我们将对抗评论分为事实扭曲、逻辑混淆和情绪操纵三类,并利用大语言模型生成具有类别特异性的扰动。框架核心是信息狄利克雷重采样(IDR)机制,可动态调整训练中恶意评论的比例,引导优化聚焦于模型最脆弱区域。实验表明,该方法在三个基准数据集上均达到领先性能,F1分数分别提升17.9%、14.5%和9.0%。
原文摘要 · Abstract (English)
Online fake news profoundly distorts public judgment and erodes trust in social platforms. While existing detectors achieve competitive performance on benchmark datasets, they remain notably vulnerable to malicious comments designed specifically to induce misclassification. This evolving threat landscape necessitates detection systems that simultaneously prioritize predictive accuracy and structural robustness. However, current detectors often fail to generalize across diverse and novel comment attack patterns. To bridge this gap, we propose AdComment, an adaptive adversarial training framework for robustness enhancement against diverse malicious comments. Based on cognitive psychology, we categorize adversarial comments into Fact Distortion, Logical Confusion, and Emotional Manipulation, and leverage LLMs to synthesize diverse, category-specific perturbations. Central to our framework is an InfoDirichlet Resampling (IDR) mechanism that dynamically adjusts malicious comment proportions during training, thereby steering optimization toward the model's most susceptible regions. Experimental results demonstrate that our approach achieves state-of-the-art performance on three benchmark datasets, improving the F1 scores by 17.9%, 14.5% and 9.0%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。