arXiv:2502.16550cs.CL2025-02EMNLP被引 8

用大模型实现多语言假新闻检测与解释生成

PropXplain: Can LLMs Enable Explainable Propaganda Detection?

  • 构建首个支持阿拉伯语和英语的带解释标注数据集
  • 模型在检测准确率相当的前提下可生成理由性解释
  • 适合需要透明化内容审核的AI研究者使用

针对多模态、多语言传播内容检测中缺乏解释的问题,本文提出首个支持阿拉伯语和英语的带解释标注数据集。同时引入一种增强解释能力的大语言模型,可在保持检测性能的同时生成基于理由的解释。实验表明该模型在检测效果相近的情况下能有效生成解释性内容。相关数据集与代码已开源(https://github.com/firojalam/PropXplain),供学术界使用。

原文摘要 · Abstract (English)

There has been significant research on propagandistic content detection across different modalities and languages. However, most studies have primarily focused on detection, with little attention given to explanations justifying the predicted label. This is largely due to the lack of resources that provide explanations alongside annotated labels. To address this issue, we propose a multilingual (i.e., Arabic and English) explanation-enhanced dataset, the first of its kind. Additionally, we introduce an explanation-enhanced LLM for both label detection and rationale-based explanation generation. Our findings indicate that the model performs comparably while also generating explanations. We will make the dataset and experimental resources publicly available for the research community (https://github.com/firojalam/PropXplain).

假新闻检测大模型可解释性多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。