arXiv:2608.21315cs.CL2026-08

揭示提示与模型交互的固定点结构,发现其不依赖任务且难以被现有机制解释。

Prompt-Model Interaction Reaches the Fixed Points: A deterministic, task-free structural readout -- and the factorizations of it that failed

  • 通过无任务的确定性读出,研究提示-模型交互在短窗口内的固定点行为。
  • 九个词的提示可完全改变固定点分布,影响模型排序,而指令微调无效。
  • 现有理论无法解释该现象,需以提示-模型对为单位进行分析。

提示的效果并非提示本身的属性:针对某一模型优化的提示在另一模型上性能下降,且中性重格式化后排名会重排。这种证据基于任务准确率,无法判断交互是任务机制问题还是条件分布本身所致。我们考察一个无任务的读出:短窗口下 argmax 映射的固定点结构,其中 x_{t+1} = argmax_x p(x | x_{t-1}, x_t),从96个起始点采样。该结构是确定性的,仅存在于短窗口——六种模型中有四种在窗口16时完全丧失。两个核心发现:第一,交互以全幅值抵达该读出:九个词的条件使固定点比例跨越大部分范围,改变四分类结构,重排模型顺序;而指令微调提升60.5 IFEval点却未改变结构。第二,所有提出机制均失效:前缀长度非单调,四种现象学因素(文体/标记、通用方向、双向性、指令抗性)在样本扩大后即消失。最接近的机制解释——早期词的注意力主导——仅在2/5模型上预测正确(随机水平),且在真实文本有效而在均匀随机输入上失败,说明当前处于其适用范围外。一个固定的九词前缀使四模型趋向0,两模型趋向1;双向性在分布内起始中仍存。此读出下,解释单元是提示-模型对。我们曾反复犯的错误有名称:将具有形状的标准施用于无变化空间的量。

原文摘要 · Abstract (English)

That a prompt's effect is not a property of the prompt is established: prompts optimised for one model degrade on another, and rankings reorder under neutral reformatting. That evidence is about task accuracy, which cannot say whether the interaction is a fact about task machinery or about the conditional distribution itself. We ask on a readout with no task in it: the fixed-point structure of the short-window argmax map x_{t+1} = argmax_x p(x | x_{t-1}, x_t), censused from 96 starts. It is deterministic, so nothing can be helped or hurt, and it exists only at short windows -- four of six models lose it entirely by window 16 -- so everything here concerns how a model reads a fragment. Two results. First, the interaction reaches this readout at full magnitude: nine tokens of conditioning move the fixed-point fraction across most of its range, change a four-way structural class, and reorder models, while instruction tuning worth 60.5 IFEval points moves the class by zero. Second, nothing we proposed carries it. Prefix length fails: the effect is not monotone. Four phenomenological factors -- prose-versus-markup, a universal direction, bidirectionality, instruct-resistance -- were each withdrawn within one run of being proposed, dissolved by widening the sample. And the nearest mechanistic account, attention-sink dominance of early tokens, predicts the sign of the shift on 2 of 5 models -- chance -- while a length-by-content cross shows it holds on real text and fails on our probe's uniformly random input, so we are outside its regime, not against it. One fixed nine-token prefix drives four models toward 0 and two toward 1; the bidirectionality survives in-distribution starts. On this readout the unit of explanation is the prompt-model pair. The recurring error it caught in us has a name: a criterion with a shape applied to a quantity with no room to vary.

提示工程模型行为固定点交互机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。