研究大模型自指时的矩阵动态,发现非闭合真值递归导致显著不稳定性。
When Self-Reference Fails to Close: Matrix-Level Dynamics in Large Language Models
- 通过多轮分析与106项指标,识别出非闭合真值递归是核心不稳源。
- 在70B模型中,注意力有效秩等指标效应量达d=3.14,显著高于稳定自指。
- 结果对理解模型自我认知失败有实际意义,适合关注模型内部机制的研究者。
我们研究自指输入如何改变大语言模型的内部矩阵动态。在四个来自三种架构家族的模型(Qwen3-VL-8B、Llama-3.2-11B、Llama-3.3-70B、Gemma-2-9B)上,对300个提示进行最多14层的14级层次结构分析,覆盖三个温度(T ∈ {0.0, 0.3, 0.7}),测量106个标量指标。发现仅自指不会引发不稳定:具象化自指与元认知提示在关键崩溃相关指标上明显比悖论性自指更稳定,甚至可与事实控制组相当。不稳定性集中于引发非闭合真值递归(NCTR)的提示——即无法在有限深度内求解的真值计算。这类提示导致注意力有效秩异常升高(表明注意力全局分散而非简单集中坍缩),在70B模型中,注意力有效秩效应量达Cohen's d=3.14,方差峰度达d=3.52;397个指标-模型组合中有281个经FDR校正后差异显著(q < 0.05),其中198个|d| > 0.8。逐层SVD验证所有采样层均出现破坏(三个模型中d > +1.0),排除聚合伪影。分类器实现AUC 0.81–0.90;30组最小差异对中42/387个组合显著;106项指标中有43项在全部四模型中复现。我们将这些现象关联至三个经典矩阵半群问题,并提出猜想:NCTR迫使有限深度Transformer趋向于这些问题集中的动力学状态。此外,NCTR提示还导致输出矛盾率上升34–56个百分点,提示其在实践中的重要性。
原文摘要 · Abstract (English)
We investigate how self-referential inputs alter the internal matrix dynamics of large language models. Measuring 106 scalar metrics across up to 7 analysis passes on four models from three architecture families -- Qwen3-VL-8B, Llama-3.2-11B, Llama-3.3-70B, and Gemma-2-9B -- over 300 prompts in a 14-level hierarchy at three temperatures ($T \in \{0.0, 0.3, 0.7\}$), we find that self-reference alone is not destabilizing: grounded self-referential statements and meta-cognitive prompts are markedly more stable than paradoxical self-reference on key collapse-related metrics, and on several such metrics can be as stable as factual controls. Instability concentrates in prompts inducing non-closing truth recursion (NCTR) -- truth-value computations with no finite-depth resolution. NCTR prompts produce anomalously elevated attention effective rank -- indicating attention reorganization with global dispersion rather than simple concentration collapse -- and key metrics reach Cohen's $d = 3.14$ (attention effective rank) to $3.52$ (variance kurtosis) vs. stable self-reference in the 70B model; 281/397 metric-model combinations differentiate NCTR from stable self-reference after FDR correction ($q < 0.05$), 198 with $|d| > 0.8$. Per-layer SVD confirms disruption at every sampled layer ($d > +1.0$ in all three models analyzed), ruling out aggregation artifacts. A classifier achieves AUC $0.81$-$0.90$; 30 minimal pairs yield 42/387 significant combinations; 43/106 metrics replicate across all four models. We connect these observations to three classical matrix-semigroup problems and propose, as a conjecture, that NCTR forces finite-depth transformers toward dynamical regimes where these problems concentrate. NCTR prompts also produce elevated contradictory output ($+34$-$56$ percentage points vs. controls), suggesting practical relevance for understanding self-referential failure modes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。