用社交推理游戏测试大模型欺骗能力,发现其比真人更难被识破。
Hidden in Plain Text: Measuring LLM Deception Quality Against Human Baselines Using Social Deduction Games
- 在异步多智能体框架中模拟35场狼人杀,用GPT-4o扮演角色
- 大模型对局的杀手识别准确率低于真人对局,说明欺骗更自然
- 适合关注AI安全、社会推理与对抗性评估的研究者
大型语言模型(LLM)代理在各类应用中日益普及,引发对其安全性的担忧。尽管已有研究证明LLM能在受控任务中实施欺骗,但其在自然语言社交情境下的欺骗能力仍不明确。本文通过社交推理游戏《狼人杀》(Mafia)研究该问题,其中成功依赖于通过对话欺骗他人。与以往研究不同,本文采用异步多智能体框架,更贴近真实社交场景。我们使用GPT-4o模拟了35场狼人杀游戏,并利用GPT-4-Turbo构建一个杀手检测器,基于游戏对话记录预测杀手身份,但不提供角色信息。以检测准确率为欺骗质量的代理指标。将该准确率与28场真人对局及随机基线对比,结果显示:在所有游戏天数和检测出的杀手数量下,模型对局的检测准确率均显著低于真人对局。这表明大模型在社交互动中更善于伪装,欺骗效果更优。我们还公开发布了一组大模型狼人杀对话语料,以支持后续研究。结果凸显了大模型在社交语境中欺骗行为的复杂性及其潜在风险。
原文摘要 · Abstract (English)
Large Language Model (LLM) agents are increasingly used in many applications, raising concerns about their safety. While previous work has shown that LLMs can deceive in controlled tasks, less is known about their ability to deceive using natural language in social contexts. In this paper, we study deception in the Social Deduction Game (SDG) Mafia, where success is dependent on deceiving others through conversation. Unlike previous SDG studies, we use an asynchronous multi-agent framework which better simulates realistic social contexts. We simulate 35 Mafia games with GPT-4o LLM agents. We then create a Mafia Detector using GPT-4-Turbo to analyze game transcripts without player role information to predict the mafia players. We use prediction accuracy as a surrogate marker for deception quality. We compare this prediction accuracy to that of 28 human games and a random baseline. Results show that the Mafia Detector's mafia prediction accuracy is lower on LLM games than on human games. The result is consistent regardless of the game days and the number of mafias detected. This indicates that LLMs blend in better and thus deceive more effectively. We also release a dataset of LLM Mafia transcripts to support future research. Our findings underscore both the sophistication and risks of LLM deception in social contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。