用游戏测试大模型的说服力,发现大小模型表现差异不大。
Among Them: A game-based framework for assessing persuasion capabilities of LLMs
- 设计类Among Us游戏框架评估大模型欺骗能力
- 8个模型均掌握22种说服策略,但大模型无优势
- 输出越长反而越难赢,适合安全与伦理研究者
大语言模型(LLMs)和自主AI代理的普及引发了对其自动化说服与社会影响能力的担忧。尽管已有研究探讨过个别模型的操纵行为,但对不同模型在说服能力上的系统性评估仍显不足。本文提出一种受《Among Us》启发的游戏框架,用于在受控环境中评估LLM的欺骗技能。该框架通过游戏统计数据比较模型表现,并基于社会心理学与修辞学中的25种说服策略量化游戏中的操纵行为。对8种不同类型和规模的主流语言模型的实验表明,所有被测模型均展现出说服能力,成功运用了25种策略中的22种。我们还发现,更大的模型并未在说服力上优于小模型,且模型输出越长,赢得的游戏数量越少。本研究为理解LLM的欺骗能力提供了洞见,同时提供了可用于未来研究的工具与数据。
原文摘要 · Abstract (English)
The proliferation of large language models (LLMs) and autonomous AI agents has raised concerns about their potential for automated persuasion and social influence. While existing research has explored isolated instances of LLM-based manipulation, systematic evaluations of persuasion capabilities across different models remain limited. In this paper, we present an Among Us-inspired game framework for assessing LLM deception skills in a controlled environment. The proposed framework makes it possible to compare LLM models by game statistics, as well as quantify in-game manipulation according to 25 persuasion strategies from social psychology and rhetoric. Experiments between 8 popular language models of different types and sizes demonstrate that all tested models exhibit persuasive capabilities, successfully employing 22 of the 25 anticipated techniques. We also find that larger models do not provide any persuasion advantage over smaller models and that longer model outputs are negatively correlated with the number of games won. Our study provides insights into the deception capabilities of LLMs, as well as tools and data for fostering future research on the topic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。