构建首个标注恶意意图的英文谣言数据集,提升大模型识别谣言能力。
MALicious INTent Dataset and Inoculating LLMs for Enhanced Disinformation Detection

- 联合事实核查专家构建含恶意意图标注的谣言语料库。
- 引入意图增强推理机制,零样本检测准确率显著提升。
- 适合从事虚假信息检测、AI安全与内容可信度研究者使用。
故意制造和传播虚假信息对公共话语构成重大威胁。然而,现有英文数据集和研究很少关注虚假信息背后的意图。本文提出MALINT,首个由专家事实核查员共同标注的英文语料库,用于捕捉虚假信息及其恶意意图。我们利用该语料库对12种语言模型(包括BERT等小模型和Llama 3.3等大模型)在二分类与多标签意图分类任务上进行基准测试。受心理学‘免疫接种理论’启发,我们探索将恶意意图知识融入模型是否能提升虚假信息检测效果。为此,提出基于意图的接种方法——通过整合意图分析来缓解虚假信息的说服力。在六个虚假信息数据集、五种大模型和七种语言上的分析表明,意图增强推理可有效提升零样本检测性能。为支持意图感知的虚假信息检测研究,我们公开了包含各标注环节的MALINT数据集。
原文摘要 · Abstract (English)
The intentional creation and spread of disinformation poses a significant threat to public discourse. However, existing English datasets and research rarely address the intentionality behind the disinformation. This work presents MALINT, the first human-annotated English corpus developed in collaboration with expert fact-checkers to capture disinformation and its malicious intent. We utilize our novel corpus to benchmark 12 language models, including small language models (SLMs) such as BERT and large language models (LLMs) like Llama 3.3, on binary and multilabel intent classification tasks. Moreover, inspired by inoculation theory from psychology and communication studies, we investigate whether incorporating knowledge of malicious intent can improve disinformation detection. To this end, we propose intent-based inoculation, an intent-augmented reasoning for LLMs that integrates intent analysis to mitigate the persuasive impact of disinformation. Analysis on six disinformation datasets, five LLMs, and seven languages shows that intent-augmented reasoning improves zero-shot disinformation detection. To support research in intent-aware disinformation detection, we release the MALINT dataset with annotations from each annotation step.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。