揭秘机器生成提示词的内在逻辑,发现其并非完全不可理解。
Evil twins are not that evil: Qualitative insights into machine-generated prompts
- 分析6种不同规模语言模型的机器生成提示,发现末尾词最关键。
- 大部分前序词可删减,仅少数关键词与生成内容有松散语义关联。
- 人类专家能准确识别关键提示词,说明提示词并非完全黑箱。
已有广泛观察表明,语言模型(LMs)对看似无意义的算法生成提示会做出可预测的响应。这既反映了我们对语言模型工作机制的理解不足,也带来了实际挑战,因为这种不透明性可能被用于恶意用途,如越狱攻击。本文首次对6种不同规模和架构的语言模型的机动生成提示(autoprompts)进行了全面分析。研究发现,这些提示的最后一个词往往具有可理解性,并显著影响生成结果;部分前置词可被删除,可能是优化过程中固定令牌数量的副产物。其余词可分为两类:填充词,可替换为无意义内容;关键词,虽无完整语法关系,但与生成内容存在松散语义关联。此外,人类专家可可靠地事后识别出最具影响力的提示词,表明这些提示并非完全不透明。最后,部分对提示的消融实验在自然语言输入中也产生类似效果,暗示这类提示是语言模型处理语言输入时的自然产物。
原文摘要 · Abstract (English)
It has been widely observed that language models (LMs) respond in predictable ways to algorithmically generated prompts that are seemingly unintelligible. This is both a sign that we lack a full understanding of how LMs work, and a practical challenge, because opaqueness can be exploited for harmful uses of LMs, such as jailbreaking. We present the first thorough analysis of opaque machine-generated prompts, or autoprompts, pertaining to 6 LMs of different sizes and families. We find that machine-generated prompts are characterized by a last token that is often intelligible and strongly affects the generation. A small but consistent proportion of the previous tokens are prunable, probably appearing in the prompt as a by-product of the fact that the optimization process fixes the number of tokens. The remaining tokens fall into two categories: filler tokens, which can be replaced with semantically unrelated substitutes, and keywords, that tend to have at least a loose semantic relation with the generation, although they do not engage in well-formed syntactic relations with it. Additionally, human experts can reliably identify the most influential tokens in an autoprompt a posteriori, suggesting these prompts are not entirely opaque. Finally, some of the ablations we applied to autoprompts yield similar effects in natural language inputs, suggesting that autoprompts emerge naturally from the way LMs process linguistic inputs in general.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。