大模型检测宣传手法能力有限,仍不如传统模型。
Are Large Language Models Good at Detecting Propaganda?
- 对比GPT系列与RoBERTa-CRF等模型在宣传手法识别中的表现。
- GPT-4 F1仅0.16,远低于RoBERTa-CRF的0.67。
- 仅在三种宣传技巧上优于基线,整体表现不突出。
宣传者常使用逻辑谬误和情感煽动等修辞手段达成目的,识别这些技巧对做出理性判断至关重要。近年来自然语言处理技术的发展使识别操纵性内容成为可能。本研究评估了几种大型语言模型在新闻文章中识别宣传手法的表现,并与基于Transformer的模型进行比较。结果发现,尽管GPT-4的F1分数(0.16)高于GPT-3.5和Claude 3 Opus,但显著低于RoBERTa-CRF基线(F1=0.67)。此外,所有三个LLM在六种宣传手法之一(人身攻击)的检测上优于多粒度网络(MGN)基线,且GPT-3.5与GPT-4在‘恐惧煽动’和‘旗帜挥舞’两种手法上也表现更优。
原文摘要 · Abstract (English)
Propagandists use rhetorical devices that rely on logical fallacies and emotional appeals to advance their agendas. Recognizing these techniques is key to making informed decisions. Recent advances in Natural Language Processing (NLP) have enabled the development of systems capable of detecting manipulative content. In this study, we look at several Large Language Models and their performance in detecting propaganda techniques in news articles. We compare the performance of these LLMs with transformer-based models. We find that, while GPT-4 demonstrates superior F1 scores (F1=0.16) compared to GPT-3.5 and Claude 3 Opus, it does not outperform a RoBERTa-CRF baseline (F1=0.67). Additionally, we find that all three LLMs outperform a MultiGranularity Network (MGN) baseline in detecting instances of one out of six propaganda techniques (name-calling), with GPT-3.5 and GPT-4 also outperforming the MGN baseline in detecting instances of appeal to fear and flag-waving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。