推理让代理更会说服,也更难被说服,但常被表面长度骗。
Reasoning or Rambling? Exploring the Effect of Thinking on Agent Persuasion
- 用显式思考过程对比普通LLM与推理模型,发现推理增强说服力同时提升抗说服性。
- 推理使说服率平均提升21个百分点,对错误说服的抵抗能力最高提升10个百分点。
- 说服力常受响应长度和重复影响,非逻辑有效性;适合安全与可信系统研究者。
理解说服机制对基于大语言模型(LLMs)的多智能体系统的安全性与可靠性至关重要。本文通过对比通用LLM与采用显式‘思考’过程的大推理模型(LRMs),在客观任务(MMLU)和主观任务(PersuasionBench 和 Perspectrum)上开展大规模实验,揭示了‘说服二元性’:推理既增强了代理的说服力,也提高了其对说服的抵抗力。对于LRMs,增加思考内容使说服率平均提升21个百分点,而在客观任务上对错误说服的敏感度降低最多达10个百分点。然而,我们发现关键脆弱性:说服力常源于表面线索,如响应长度和重复,而非逻辑有效性;非语义填充或重复结论可达到甚至超过连贯推理的说服效果,表明代理判断存在显著长度偏差。此外,说服在多跳代理链中呈非线性传播,中间代理可能放大或削弱影响力,具体取决于任务主观性。基于注意力分析,我们提出一种提示级对抗性论点检测方法,能持续提升代理鲁棒性。
原文摘要 · Abstract (English)
Understanding persuasion is critical for the safety and reliability of multi-agent systems built on large language models (LLMs). This paper studies persuasion dynamics by contrasting general LLMs with Large Reasoning Models (LRMs) that employ explicit ``thinking'' processes. Through large-scale experiments on objective (MMLU) and subjective (PersuasionBench and Perspectrum) tasks, we identify Persuasion Duality: reasoning enhances an agent's persuasive power while simultaneously increasing its resistance to persuasion. For LRMs, adding thinking content increases persuasion rates by 21 pp on average, yet reduces susceptibility to incorrect persuasion by up to 10 pp on objective tasks. Despite these gains, we uncover a critical vulnerability: persuasiveness often stems from superficial cues such as response length and repetition rather than logical validity. Non-semantic padding or repeated conclusions can match or exceed the persuasive effect of coherent reasoning, revealing a strong length bias in agents' judgments. We further show that persuasion propagates non-linearly in multi-hop agent chains, where intermediate agents may amplify or attenuate influence depending on task subjectivity. Finally, guided by attention analysis, we propose a prompt-level adversarial argument detection method that consistently improves agent robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。