测试大模型在保密前提下精准传递信息的能力。
SNEAK: Evaluating Strategic Communication and Information Leakage in Large Language Models
- 设计双角色模拟:盟友与伪装者,评估信息传递与泄露程度。
- 人类表现远超模型,最高得分是模型的四倍。
- 揭示当前大模型在策略性沟通上仍有明显短板。
大型语言模型(LLMs)越来越多地部署于多智能体场景中,需在信息传递与保密之间取得平衡。在该场景中,智能体可能需向合作者透露信息,同时防止对手推断敏感内容。然而,现有评测基准主要关注推理、事实知识或指令遵循能力,未能直接衡量在信息不对称下的战略沟通。本文提出SNEAK(面向对抗性知识的秘密感知自然语言评测),用于评估语言模型在选择性信息共享中的表现。在SNEAK任务中,模型需根据一个语义类别、一组候选词及一个秘密词,生成一条既表明知晓秘密又不直接暴露秘密的信息。通过两个模拟代理——知情盟友(需识别意图)与不知情伪装者(试图推断秘密)——来评估生成消息的效用与泄漏程度。该框架产生两项互补指标:效用(对合作者的传达效果)与泄漏(对对手的信息暴露)。基于此,我们分析了现代语言模型在信息不对称下的权衡表现,发现当前系统在战略沟通方面仍具挑战性。值得注意的是,人类参与者显著优于所有被测模型,最高得分达模型的四倍。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in multi-agent settings where communication must balance informativeness and secrecy. In such settings, an agent may need to signal information to collaborators while preventing an adversary from inferring sensitive details. However, existing LLM benchmarks primarily evaluate capabilities such as reasoning, factual knowledge, or instruction following, and do not directly measure strategic communication under asymmetric information. We introduce SNEAK (Secret-aware Natural language Evaluation for Adversarial Knowledge), a benchmark for evaluating selective information sharing in language models. In SNEAK, a model is given a semantic category, a candidate set of words, and a secret word, and must generate a message that indicates knowledge of the secret without revealing it too clearly. We evaluate generated messages using two simulated agents with different information states: an ally, who knows the secret and must identify the intended message, and a chameleon, who does not know the secret and attempts to infer it from the message. This yields two complementary metrics: utility, measuring how well the message communicates to collaborators, and leakage, measuring how much information it reveals to an adversary. Using this framework, we analyze the trade-off between informativeness and secrecy in modern language models and show that strategic communication under asymmetric information remains a challenging capability for current systems. Notably, human participants outperform all evaluated models by a large margin, achieving up to four times higher scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。