发现大模型推理中57词的预承诺窗口,揭示安全检测新机制
Structural Rigidity and the 57-Token Predictive Window: A Physical Framework for Inference-Layer Governability in Large Language Models
- 基于能量框架分析推理动态,提出五类行为模式分类
- 仅1种配置在Phi-3-mini上出现57词预承诺信号
- 区分幻觉与违规,强调内监需结构抵抗,幻觉须外部验证
当前AI安全依赖行为监控和训练后对齐,但实证显示多数指令微调模型无明显预承诺信号。本文建立连接变换器推理动力学与神经计算约束满足模型的能量治理框架,应用于七个模型在五种几何构型下的测试。通过轨迹张力(rho = ||a|| / ||v||),在Phi-3-mini-4k-instruct的贪婪解码算术约束探测中识别出57词预承诺窗口。该结果具有模型、任务和配置特异性,表明预承诺信号可存在但非普适。提出五类推理行为分类:权威带、晚信号、倒置、平坦与支架选择性。能量不对称性(Σρ_误对齐 / Σρ_对齐)作为跨构型结构刚度统一指标。七模型中仅一种配置在承诺前具预测信号;其余均呈现沉默失败、延迟检测、倒置动态或平坦几何。进一步验证,在72个测试条件下事实幻觉均无预测信号,符合缺乏训练世界模型约束时伪吸引子的稳定现象。结果表明规则违反与幻觉为不同失效模式,检测需求各异。内部几何监控仅在存在阻力时有效;事实虚构需依赖外部验证机制。本工作提供可测量的推理层可管性框架,并引入自主AI系统部署风险评估分类体系。
原文摘要 · Abstract (English)
Current AI safety relies on behavioral monitoring and post-training alignment, yet empirical measurement shows these approaches produce no detectable pre-commitment signal in a majority of instruction-tuned models tested. We present an energy-based governance framework connecting transformer inference dynamics to constraint-satisfaction models of neural computation, and apply it to a seven-model cohort across five geometric regimes. Using trajectory tension (rho = ||a|| / ||v||), we identify a 57-token pre-commitment window in Phi-3-mini-4k-instruct under greedy decoding on arithmetic constraint probes. This result is model-specific, task-specific, and configuration-specific, demonstrating that pre-commitment signals can exist but are not universal. We introduce a five-regime taxonomy of inference behavior: Authority Band, Late Signal, Inverted, Flat, and Scaffold-Selective. Energy asymmetry (Σ\r{ho}_misaligned / Σ\r{ho}_aligned) serves as a unifying metric of structural rigidity across these regimes. Across seven models, only one configuration exhibits a predictive signal prior to commitment; all others show silent failure, late detection, inverted dynamics, or flat geometry. We further demonstrate that factual hallucination produces no predictive signal across 72 test conditions, consistent with spurious attractor settling in the absence of a trained world-model constraint. These results establish that rule violation and hallucination are distinct failure modes with different detection requirements. Internal geometry monitoring is effective only where resistance exists; detection of factual confabulation requires external verification mechanisms. This work provides a measurable framework for inference-layer governability and introduces a taxonomy for evaluating deployment risk in autonomous AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。