arXiv:2509.22989cs.AIcs.CY2025-09被引 11

用贝叶斯理论框架评估并训练大模型做策略性说服

Towards Strategic Persuasion with Language Models

  • 基于贝叶斯说服理论,重构人类说服数据集构建评估环境
  • 前沿模型在多场景下实现显著说服增益,策略符合理论预期
  • 小模型经强化学习训练后说服效果大幅提升,具可扩展性

大语言模型(LLMs)展现出与人类相当的说服能力,带来潜在价值的同时也引发社会担忧。然而,系统评估其说服能力极具挑战,因人类说服效果在不同领域差异显著。本文采用理论驱动方法,提出一个可扩展、有原则的框架,研究LLMs的说服能力。基于贝叶斯说服理论,我们重构人类-人类说服数据集,构建用于评估和训练LLMs作为策略性说服者的环境。结果表明,前沿模型能持续获得高说服增益,并表现出与理论一致的复杂说服策略。在此基础上,我们使用强化学习训练LLMs进行策略性说服。结果显示,即使小型模型通过强化学习也能显著提升说服效果。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated strong persuasive capabilities comparable to those of humans, offering promising benefits while raising societal concerns. However, systematically evaluating the persuasive capabilities of LLMs is inherently challenging, as the effectiveness of persuasion among humans varies significantly across different domains. In this paper, we take a theory-driven approach to provide a scalable and principled framework for studying the persuasive capabilities of LLMs. Grounded in Bayesian persuasion theory, we repurpose human-human persuasion datasets to construct environments for evaluating and training LLMs as strategic persuaders. Our results reveal that frontier models can consistently achieve high persuasion gains and exhibit sophisticated persuasion strategies that align with theoretical characterizations. Building on this, we use reinforcement learning to train LLMs for strategic persuasion in our environments. Our results also demonstrate that even small LLMs can obtain significantly higher persuasion gains through reinforcement learning.

语言模型策略说服强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。