arXiv:2603.09884cs.CLcs.CY2026-03被引 1

测试前沿大模型政治说服力,发现部分模型比广告更有效。

Benchmarking Political Persuasion Risks Across Frontier Large Language Models

  • 用问卷实验对比七款主流大模型在政治议题上的说服效果。
  • Claude模型最能说服人,Grok最弱,结果跨议题稳定。
  • 信息型提示对不同模型效果相反,适合做风险评估参考。

大型语言模型(LLM)在政治观点影响方面的能力引发持续担忧。尽管以往研究认为LLM的说服力不高于传统政治宣传,但前沿模型的兴起需要重新评估。我们在跨党派议题和立场上开展了两项调查实验(N=19,145),评估了Anthropic、OpenAI、Google和xAI开发的七款先进LLM。结果表明,LLM整体优于标准竞选广告,模型间表现存在差异:Claude系列最具说服力,Grok最低,且结果在不同议题与立场下均稳健。此外,与Hackenburg等(2025b)和Lin等(2025)发现信息提示提升说服力不同,我们发现信息提示的效果具有模型依赖性——它增强Claude和Grok的说服力,却显著降低GPT的说服效果。本文提出一种数据驱动、策略无关的LLM辅助对话分析方法,用于识别和评估潜在说服策略。研究为前沿模型的说服风险提供基准,并建立跨模型比较评估框架。

原文摘要 · Abstract (English)

Concerns persist regarding the capacity of Large Language Models (LLMs) to sway political views. Although prior research has claimed that LLMs are not more persuasive than standard political campaign practices, the recent rise of frontier models warrants further study. In two survey experiments (N=19,145) across bipartisan issues and stances, we evaluate seven state-of-the-art LLMs developed by Anthropic, OpenAI, Google, and xAI. We find that LLMs outperform standard campaign advertisements, with heterogeneity in performance across models. Specifically, Claude models exhibit the highest persuasiveness, while Grok exhibits the lowest. The results are robust across issues and stances. Moreover, in contrast to the findings in Hackenburg et al. (2025b) and Lin et al. (2025) that information-based prompts boost persuasiveness, we find that the effectiveness of information-based prompts is model-dependent: they increase the persuasiveness of Claude and Grok while substantially reducing that of GPT. We introduce a data-driven and strategy-agnostic LLM-assisted conversation analysis approach to identify and assess underlying persuasive strategies. Our work benchmarks the persuasive risks of frontier models and provides a framework for cross-model comparative risk assessment.

大模型风险政治说服评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。