用大模型分析特朗普演讲中的精细型民粹主义,发现小模型微调后更准确。
Identifying Fine-grained Forms of Populism in Political Discourse: A Case Study on Donald Trump's Presidential Campaigns
- 构建专用数据集,测试大模型识别民粹话语的能力。
- 微调的RoBERTa模型远超主流指令微调大模型,除非也微调。
- 首次在跨政治语境下验证模型迁移性,适合政治话语研究者。
大语言模型(LLMs)在多种指令跟随任务中表现出色,但在理解社会科学研究中的细微概念方面仍不充分。本文探讨了LLMs是否能识别并分类精细形态的民粹主义——这一在学术与媒体争论中复杂且争议不断的概念。为此,我们构建并发布了专为捕捉民粹话语而设计的新数据集。评估了多种预训练语言模型(包括开源与闭源)在不同提示范式下的表现。结果揭示性能显著差异,凸显了大模型在识别民粹话语方面的局限性。发现微调后的RoBERTa分类器远超所有新世代指令微调大模型,除非同样进行微调。进一步将最优模型应用于唐纳德·特朗普总统竞选演讲,提取其策略性运用民粹修辞的关键洞见。最后,通过在欧洲政客竞选演讲上进行基准测试,评估模型的泛化能力,发现指令微调大模型在域外数据上更具鲁棒性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of instruction-following tasks, yet their grasp of nuanced social science concepts remains underexplored. This paper examines whether LLMs can identify and classify fine-grained forms of populism, a complex and contested concept in both academic and media debates. To this end, we curate and release novel datasets specifically designed to capture populist discourse. We evaluate a range of pre-trained (large) language models, both open-weight and proprietary, across multiple prompting paradigms. Our analysis reveals notable variation in performance, highlighting the limitations of LLMs in detecting populist discourse. We find that a fine-tuned RoBERTa classifier vastly outperforms all new-era instruction-tuned LLMs, unless fine-tuned. Additionally, we apply our best-performing model to analyze campaign speeches by Donald Trump, extracting valuable insights into his strategic use of populist rhetoric. Finally, we assess the generalizability of these models by benchmarking them on campaign speeches by European politicians, offering a lens into cross-context transferability in political discourse analysis. In this setting, we find that instruction-tuned LLMs exhibit greater robustness on out-of-domain data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。