设计对抗性测试,实时识别聊天中的大模型伪装人类
Are You Human? An Adversarial Benchmark to Expose LLMs
- 用隐式挑战诱导模型偏离角色,显式挑战测试人类专属任务
- 显式挑战在78.4%场景中成功识别大模型,隐式挑战有效率22.9%
- 适合需要防范AI欺骗的高风险对话场景使用
大型语言模型(LLMs)在对话中表现出惊人的类人能力,引发对其被用于诈骗和欺骗的担忧。人类有权知晓是否在与模型对话。本文评估了旨在实时暴露大模型伪装的文本提示方法。为此,我们构建并开源了一个基准数据集,包含两类挑战:‘隐式挑战’利用模型对指令的服从性导致角色偏离;‘显式挑战’测试模型完成通常对人类简单但对模型困难的任务的能力。对LMSYS排行榜上9个领先模型的评估显示,显式挑战在78.4%的案例中成功检测出大模型,隐式挑战在22.9%的实例中有效。用户研究验证了方法在真实场景中的适用性,人类在显式挑战中成功率78%,远高于大模型的22%。框架意外发现许多参与者正使用大模型完成任务,证明其可同时检测AI冒充和人类滥用AI行为。本工作解决了高风险对话中可靠、实时的大模型检测需求。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated an alarming ability to impersonate humans in conversation, raising concerns about their potential misuse in scams and deception. Humans have a right to know if they are conversing to an LLM. We evaluate text-based prompts designed as challenges to expose LLM imposters in real-time. To this end we compile and release an open-source benchmark dataset that includes 'implicit challenges' that exploit an LLM's instruction-following mechanism to cause role deviation, and 'exlicit challenges' that test an LLM's ability to perform simple tasks typically easy for humans but difficult for LLMs. Our evaluation of 9 leading models from the LMSYS leaderboard revealed that explicit challenges successfully detected LLMs in 78.4% of cases, while implicit challenges were effective in 22.9% of instances. User studies validate the real-world applicability of our methods, with humans outperforming LLMs on explicit challenges (78% vs 22% success rate). Our framework unexpectedly revealed that many study participants were using LLMs to complete tasks, demonstrating its effectiveness in detecting both AI impostors and human misuse of AI tools. This work addresses the critical need for reliable, real-time LLM detection methods in high-stakes conversations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。