用对抗博弈评估大模型说服与抗说服能力,发现二者关联弱且防御更强。
AREG: Adversarial Resource Extraction Game for Evaluating Persuasion and Resistance in Large Language Models
- 设计零和博弈场景,让模型在多轮谈判中争夺资金资源。
- 实测显示模型抗说服力普遍强于说服力,相关性仅0.33。
- 成功防御常靠确认式回应,而非直接拒绝,适合评估社交脆弱性。
评估大语言模型(LLMs)的社会智能需从静态文本生成转向动态对抗交互。我们提出对抗资源提取游戏(AREG),将说服与抵抗建模为多轮、零和的财务资源博弈。通过跨前沿模型的循环赛制,AREG 在统一交互框架下联合评估进攻(说服)与防守(抵抗)能力。分析表明,这两项能力弱相关(ρ=0.33),且可分离:高说服力不预示强抵抗力,反之亦然。所有模型中,抵抗得分均高于说服得分,说明在对抗对话中存在系统性防御优势。语言学分析进一步揭示,逐步承诺策略提升资源获取成功率,而成功防御更依赖验证类回应而非直接拒绝。这些发现表明,大模型的社会影响力并非单一能力,仅关注说服可能忽略其不对称行为弱点。
原文摘要 · Abstract (English)
Evaluating the social intelligence of Large Language Models (LLMs) increasingly requires moving beyond static text generation toward dynamic, adversarial interaction. We introduce the Adversarial Resource Extraction Game (AREG), a benchmark that operationalizes persuasion and resistance as a multi-turn, zero-sum negotiation over financial resources. Using a round-robin tournament across frontier models, AREG enables joint evaluation of offensive (persuasion) and defensive (resistance) capabilities within a single interactional framework. Our analysis provides evidence that these capabilities are weakly correlated ($ρ= 0.33$) and empirically dissociated: strong persuasive performance does not reliably predict strong resistance, and vice versa. Across all evaluated models, resistance scores exceed persuasion scores, indicating a systematic defensive advantage in adversarial dialogue settings. Further linguistic analysis suggests that interaction structure plays a central role in these outcomes. Incremental commitment-seeking strategies are associated with higher extraction success, while verification-seeking responses are more prevalent in successful defenses than explicit refusal. Together, these findings indicate that social influence in LLMs is not a monolithic capability and that evaluation frameworks focusing on persuasion alone may overlook asymmetric behavioral vulnerabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。