arXiv:2503.01829cs.CLcs.AI2025-03被引 28

评测大模型的说服力与易受说服程度,发现GPT-4o最抗诱导。

Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models

  • 构建自动化多智能体对话框架,自动评估模型说服力和易受说服性。
  • GPT-4o说服力与Llama-3.3-70B相当,但对虚假信息抵抗力高50%以上。
  • o4-mini兼具强说服力与高抗性,适合安全敏感场景应用。

大型语言模型(LLMs)展现出接近人类水平的说服能力。尽管这些能力可用于社会公益,但也存在被滥用的风险。除了关注模型如何说服他人,其自身对说服的敏感性也构成关键对齐挑战,关乎鲁棒性、安全性与伦理遵循。为此,我们提出「说服我如果可能」(PMIYC)框架,一种用于评估多智能体交互中说服效果与易受说服性的自动化方法。该框架提供可扩展替代方案,避免传统依赖人工标注的高成本与耗时问题。PMIYC自动执行说服者与被说服者之间的多轮对话,量化双方的说服有效性与易感性。全面评估涵盖多种LLM及不同说服场景(如主观判断与误导信息)。通过人工评估验证框架有效性,结果与已有研究一致。实验发现:Llama-3.3-70B与GPT-4o说服力相近,优于Claude 3 Haiku达30%;而GPT-4o对误导信息的抵抗能力比Llama-3.3-70B高出50%以上。值得注意的是,o4-mini在说服他人方面表现优异,同时具备强抗诱导能力。这些发现为理解大模型的说服动态提供了实证依据,并推动更安全的AI系统发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate persuasive capabilities that rival human-level persuasion. While these capabilities can be used for social good, they also present risks of potential misuse. Beyond the concern of how LLMs persuade others, their own susceptibility to persuasion poses a critical alignment challenge, raising questions about robustness, safety, and adherence to ethical principles. To study these dynamics, we introduce Persuade Me If You Can (PMIYC), an automated framework for evaluating persuasiveness and susceptibility to persuasion in multi-agent interactions. Our framework offers a scalable alternative to the costly and time-intensive human annotation process typically used to study persuasion in LLMs. PMIYC automatically conducts multi-turn conversations between Persuader and Persuadee agents, measuring both the effectiveness of and susceptibility to persuasion. Our comprehensive evaluation spans a diverse set of LLMs and persuasion settings (e.g., subjective and misinformation scenarios). We validate the efficacy of our framework through human evaluations and demonstrate alignment with human assessments from prior studies. Through PMIYC, we find that Llama-3.3-70B and GPT-4o exhibit similar persuasive effectiveness, outperforming Claude 3 Haiku by 30%. However, GPT-4o demonstrates over 50% greater resistance to persuasion for misinformation compared to Llama-3.3-70B. Notably, o4-mini emerges as both an effective persuader, and a resistant persuadee. These findings provide empirical insights into the persuasive dynamics of LLMs and contribute to the development of safer AI systems.

大模型评估说服力对抗性测试AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。