arXiv:2505.02273cs.CL2025-05EMNLP被引 2

揭秘语言模型中优化提示的构造与作用机制

Demystifying optimized prompts in language models

  • 优化提示由罕见标点和名词构成,非自然语言
  • 模型内部激活模式显示其与正常语句明显不同
  • 跨多种指令微调模型,提示处理路径高度一致

现代语言模型对分布外输入缺乏鲁棒性。通过机器生成的‘优化提示’可调节模型输出并诱导特定行为,同时表面看似完全不可理解。本文研究优化提示的组成及其在模型中的解析机制。发现优化提示主要由训练数据中较少见的标点和名词构成;模型内部激活中,优化提示表现出与自然语言截然不同的稀疏子集。在多种指令微调模型中,优化提示的表征形成路径具有一致性。

原文摘要 · Abstract (English)

Modern language models (LMs) are not robust to out-of-distribution inputs. Machine generated (``optimized'') prompts can be used to modulate LM outputs and induce specific behaviors while appearing completely uninterpretable. In this work, we investigate the composition of optimized prompts, as well as the mechanisms by which LMs parse and build predictions from optimized prompts. We find that optimized prompts primarily consist of punctuation and noun tokens which are more rare in the training data. Internally, optimized prompts are clearly distinguishable from natural language counterparts based on sparse subsets of the model's activations. Across various families of instruction-tuned models, optimized prompts follow a similar path in how their representations form through the network.

提示工程语言模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。