arXiv:2608.10986cs.CLcs.LG2026-08被引 1

用自反馈环探测语言模型,发现测量结果混杂了构造影响与模型特性。

What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model

论文配图:What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model
图 1 · 摘自论文原文
  • 构建可重复的自反馈环,通过损伤传播度量模型动态
  • 19个模型中λ_ca在训练中跨尺度稳定过零点,具可复现性
  • 提出分离构造与模型影响的验证方法,避免误判

一类新兴方法通过将语言模型的输出回传以探查其行为:自一致性、迭代精炼、代理循环。本文设计一个精确构造的环形令牌系统,利用模型自身的窗口条件概率p_r(x_i | x_{i±r})原位重采样。该系统遵循令牌序列上的Glauber动力学,但耦合方式改变:共享随机数下两环仅差一令牌时,未受损副本完全一致,损伤传播得以量化;而最大耦合则给出混合时间。结果表明,此类探针同时测量两类性质:部分量由构造决定(如损伤光锥为运动学量,λ_ca(r)半径标度在19个模型和两个70倍尺度梯度上不变);另一部分真实反映模型(λ_ca在训练中可复现地过零,吸引子占比始终排序一致)。若不区分两者,前者易被误认为后者——作者自身曾误判四个月,报告一个精确到三位小数的相变,实属探针特性而非模型属性。本文提出分离检验:固定构造变模型,或固定模型变构造,观察读数变化。通过复现独立预测的Domany-Kinzel损伤场实现比特级一致,并识别出四个因该方法论检出而被撤回的错误结论,均源于看似可测却实为构造效应的量。

原文摘要 · Abstract (English)

A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops. We ask what such a probe measures, in a construction chosen to make the question sharp: a ring of token cells resampled in place by the model's own windowed conditional p_r(x_i | x_{i+-r}). The substrate is Glauber dynamics on token sequences and is not new; what we change is the coupling. Advancing two rings that differ in one token under common random numbers makes undamaged copies diverge by exactly zero, so damage spreading becomes measurable where a maximal coupling gives mixing times instead. The answer is that it measures two different things at once, in readings that look alike. Some quantities are fixed by the construction: the damage light cone is kinematic, and the radius scaling of the token-space Lyapunov exponent lambda_ca(r) is model-invariant across 19 models and two scale ladders spanning 70x. Others genuinely track the model: lambda_ca crosses zero at a reproducible point in training, and the attractor share ranks models consistently however the lattice is built. Left undistinguished, the first kind is readily mistaken for the second -- we did so ourselves for four months, and report a phase transition we measured to three decimal places that belongs to the probe rather than to any language model. We give the test that separates them: hold the construction fixed and vary the model, or hold the model fixed and vary the construction, and see which readings move. We validate the instrument by reproduction first, recovering a Domany-Kinzel damage field bit-exactly against an independent prediction, and we report the estimator failures that this discipline caught -- four retracted verdicts, each on a quantity that looked like a measurement. The methodology ships as a package.

语言模型自反馈动态测量探针验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。