arXiv:2604.15400cs.LGcs.AI2026-04被引 2

发现大模型幻觉是早期快速锁定的稳定状态,难以纠正。

Hallucination as Trajectory Commitment: Causal Evidence for Asymmetric Attractor Dynamics in Transformer Generation

论文配图:Hallucination as Trajectory Commitment: Causal Evidence for Asymmetric Attractor Dynamics in Transformer Generation
图 1 · 摘自论文原文
  • 通过重复输入相同提示,发现44.3%的生成路径在首个词就分叉出幻觉
  • 注入幻觉激活可使正确输出崩溃(87.5%失败),反向修复成功率仅33.3%
  • 幻觉倾向由初始提示编码决定,且可被聚类识别为五类稳定状态

我们提供了因果证据,表明自回归语言模型中的幻觉是由不对称吸引子动力学驱动的早期轨迹锁定现象。通过相同提示分叉实验,在61个跨六类主题的提示上,27个(44.3%)在第一个生成词即出现事实与幻觉路径的分离(第0步KL=0,第1步KL>1.0)。对28层激活块进行干扰测试显示:将幻觉激活注入正确路径导致87.5%试验失败(第20层),而反向操作仅能恢复33.3%(第24层),均显著高于10.4%基线和12.5%随机控制(p=0.025)。窗口扰动实验表明,纠正需多步持续干预,而破坏只需单步扰动。探测提示编码发现,第15层初始残差状态与每提示幻觉率相关性达r=0.776(p<0.001),无监督聚类识别出五个具有代表性的状态组(eta²=0.55),其中靠近鞍点的组包含13个分叉假前提提示中的12个,表明吸引子结构在提示编码阶段已由可聚类的状态决定。结果表明幻觉是局部稳定的吸引子盆地:进入快速且概率性,退出需跨层跨步协调干预,相关盆地由步骤0即可识别的可聚类状态所预选。

原文摘要 · Abstract (English)

We present causal evidence that hallucination in autoregressive language models is an early trajectory commitment governed by asymmetric attractor dynamics. Using same-prompt bifurcation, in which we repeatedly sample identical inputs to observe spontaneous divergence, we isolate trajectory dynamics from prompt-level confounds. On Qwen2.5-1.5B across 61 prompts spanning six categories, 27 prompts (44.3%) bifurcate with factual and hallucinated trajectories diverging at the first generated token (KL = 0 at step 0, KL > 1.0 at step 1). Activation patching across 28 layers reveals a pronounced causal asymmetry: injecting a hallucinated activation into a correct trajectory corrupts output in 87.5% of trials (layer 20), while the reverse recovers only 33.3% (layer 24); both exceed the 10.4% baseline (p = 0.025) and 12.5% random-patch control. Window patching shows correction requires sustained multi-step intervention, whereas corruption needs only a single perturbation. Probing the prompt encoding itself, step-0 residual states predict per-prompt hallucination rate at Pearson r = 0.776 at layer 15 (p < 0.001 against a 1000-permutation null); unsupervised clustering identifies five regime-like groups (eta^2 = 0.55) whose saddle-adjacent cluster concentrates 12 of the 13 bifurcating false-premise prompts, indicating that the basin structure is organized around regime commitments fixed at prompt encoding. These findings characterize hallucination as a locally stable attractor basin: entry is probabilistic and rapid, exit demands coordinated intervention across layers and steps, and the relevant basins are selected by clusterable regimes already discernible at step 0.

大模型幻觉生成机制吸引子动力学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。