arXiv:2409.10338cs.CLcs.AI2024-09被引 3

用20个问题区分大模型真假,效率翻倍且隐蔽性强。

The 20 questions game to distinguish large language models

  • 设计两种提问策略,通过二元问题识别模型身份
  • 仅用10个问题即可区分22个大模型,准确率接近100%
  • 适合审计或版权方检测模型泄露,隐秘性高

受二十个问题游戏启发,本文提出一种方法,在黑盒环境下判断两个大语言模型(LLMs)是否相同。目标是使用少量(通常不超过20个)良性二元问题完成判断。我们形式化该问题,并首先基于已知基准数据集随机选取问题建立基线,20个问题内准确率达近100%。在给出该任务的理论最优边界后,引入两种高效提问启发式策略,仅需一半问题即可完成对22个大模型的区分。该方法在隐蔽性上具有显著优势,适用于审计人员或版权持有者怀疑模型泄露时的检测场景。

原文摘要 · Abstract (English)

In a parallel with the 20 questions game, we present a method to determine whether two large language models (LLMs), placed in a black-box context, are the same or not. The goal is to use a small set of (benign) binary questions, typically under 20. We formalize the problem and first establish a baseline using a random selection of questions from known benchmark datasets, achieving an accuracy of nearly 100% within 20 questions. After showing optimal bounds for this problem, we introduce two effective questioning heuristics able to discriminate 22 LLMs by using half as many questions for the same task. These methods offer significant advantages in terms of stealth and are thus of interest to auditors or copyright owners facing suspicions of model leaks.

大模型鉴别黑盒测试隐私安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。