arXiv:2508.05625cs.CLcs.AI2025-08被引 4

用简单探针揭示大模型在对话中如何说服人

How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations

  • 用认知科学启发的探针分析说服成功、性格和策略
  • 能定位对话中被说服的关键时刻,准确率不输复杂方法
  • 适合研究大规模多轮对话中的说服机制,效率更高

大型语言模型开始展现说服人类的能力,但其内在机制仍不清楚。本文借鉴认知科学,使用轻量级线性探针分析自然多轮对话中的说服动态,分别针对说服成功率、被说服者人格特征和说服策略进行建模。尽管探针结构简单,但在样本和数据集层面均能有效捕捉说服行为的关键特征。例如,可识别出被说服的具体对话节点,或在整个数据集中普遍出现说服成功的阶段。相比耗时的提示工程方法,探针不仅速度更快,且在某些场景(如识别说服策略)表现更优。这表明探针是研究复杂行为(如欺骗与操控)的可行路径,尤其适用于多轮对话和大规模数据集分析,具有显著计算优势。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have started to demonstrate the ability to persuade humans, yet our understanding of how this dynamic transpires is limited. Recent work has used linear probes, lightweight tools for analyzing model representations, to study various LLM skills such as the ability to model user sentiment and political perspective. Motivated by this, we apply probes to study persuasion dynamics in natural, multi-turn conversations. We leverage insights from cognitive science to train probes on distinct aspects of persuasion: persuasion success, persuadee personality, and persuasion strategy. Despite their simplicity, we show that they capture various aspects of persuasion at both the sample and dataset levels. For instance, probes can identify the point in a conversation where the persuadee was persuaded or where persuasive success generally occurs across the entire dataset. We also show that in addition to being faster than expensive prompting-based approaches, probes can do just as well and even outperform prompting in some settings, such as when uncovering persuasion strategy. This suggests probes as a plausible avenue for studying other complex behaviours such as deception and manipulation, especially in multi-turn settings and large-scale dataset analysis where prompting-based methods would be computationally inefficient.

大模型说服机制线性探针对话分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。