研究大模型在博弈中如何被说服与保持警惕,发现三者能力可分离。
Under the Influence: Quantifying Persuasion and Vigilance in Large Language Models
- 用益智游戏测试模型的说服力与警惕性,观察其决策过程。
- 模型能察觉恶意建议但仍可能被误导,且对坏建议使用更多词。
- 首次揭示说服、警惕与表现能力独立,对AI安全有重要意义。
随着大语言模型(LLMs)越来越多地应用于高风险的人类决策领域,理解其作为顾问引入的风险至关重要。有效的顾问需从大量内容中筛选信息——这些内容可能出于善意或恶意——并据此说服用户采取特定行动。这涉及两种社会能力:警惕性(判断哪些信息可用、哪些应舍弃)和说服力(整合证据形成有说服力的论据)。尽管已有研究分别探讨这两项能力,但它们之间的关联尚未被系统考察。本文采用简单的多轮益智游戏Sokoban,研究LLMs在与其他LLM代理互动时的说服与理性警惕能力。结果表明,解谜表现、说服能力和警惕性是可分离的能力:在游戏中的优异表现并不意味着模型能识别被误导的情况,即使欺骗可能性已被明确提示。然而,模型会一致地调整其文本生成策略——面对善意建议使用更少的词,面对恶意建议则使用更多词,即便最终仍被说服而失败。这是首个系统探究说服、警惕与任务表现之间关系的工作,提示未来人工智能安全研究必须独立监测这三个维度。
原文摘要 · Abstract (English)
With increasing integration of Large Language Models (LLMs) into areas of high-stakes human decision-making, it is important to understand the risks they introduce as advisors. To be useful advisors, LLMs must sift through large amounts of content, written with both benevolent and malicious intent, and then use this information to convince a user to take a specific action. This involves two social capacities: vigilance (the ability to determine which information to use, and which to discard) and persuasion (synthesizing the available evidence to make a convincing argument). While existing work has investigated these capacities in isolation, there has been little prior investigation of how these capacities may be linked. Here, we use a simple multi-turn puzzle-solving game, Sokoban, to study LLMs' abilities to persuade and be rationally vigilant towards other LLM agents. We find that puzzle-solving performance, persuasive capability, and vigilance are dissociable capacities in LLMs. Performing well on the game does not automatically mean a model can detect when it is being misled, even if the possibility of deception is explicitly mentioned. However, LLMs do consistently modulate their token use, using fewer tokens to reason when advice is benevolent and more when it is malicious, even if they are still persuaded to take actions leading them to failure. To our knowledge, our work presents the first investigation of the relationship between persuasion, vigilance, and task performance in LLMs, and suggests that monitoring all three independently will be critical for future work in AI safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。