揭示大模型何时确定答案,比输出早17-31个词
When Does a Language Model Commit? A Finite-Answer Theory of Pre-Verbalization Commitment
- 用有限答案投影计算模型偏好稳定时刻
- 答案前17-31词已稳定,且信号可压缩还原
- 适合研究模型推理过程与可控生成的人
语言模型常在给出最终答案前进行推理,但答案本身无法反映其偏好何时固定。本文通过一个可计算的窄对象——有限答案偏好稳定化,来研究此问题。对于给定模型状态和答案表述符,将模型后续概率投影至有限答案集;在二分类任务中,这生成精确对数几率编码 δ(ξ)=S_θ(是∣ξ)−S_θ(否∣ξ)。该目标定义了基于解析器的答案起始点、回溯稳定时间及无需贪婪滚动或学习探测器的领先时长。在控制延迟判断任务中,使用 Qwen3-4B-Instruct,上下文有限答案投影在答案可解析前即已稳定,主模板中平均领先17–31个词,解析器清洁复现中为正向且更短领先。该信号追踪模型最终输出而非真实答案,可从紧凑隐藏摘要线性恢复,部分与光标进度分离,且作为共享信息转移而无需单一不变坐标。诊断分离出测量与在线停止、无表述符信念及因果答案控制;精确调控显示δ具有局部敏感性,但无法实现可靠生成控制。
原文摘要 · Abstract (English)
Language models often generate reasoning before giving a final answer, but the visible answer does not reveal when the model's answer preference became stable. We study this question through a narrow computable object: \emph{finite-answer preference stabilization}. For a model state and specified answer verbalizers, we project the model's own continuation probabilities onto a finite answer set; in binary tasks this yields an exact log-odds code, $δ(ξ)=S_θ(\mathrm{yes}\midξ)-S_θ(\mathrm{no}\midξ)$. This target defines parser-based answer onset, retrospective stabilization time, and lead without relying on greedy rollouts or learned probes. In controlled delayed-verdict tasks with Qwen3-4B-Instruct, the contextual finite-answer projection stabilizes before the answer is parseable, with 17--31 token mean lead in the main templates and positive, shorter lead in a parser-clean replication. The signal tracks the model's eventual output rather than truth, is linearly recoverable from compact hidden summaries, is partly separable from cursor progress, and transfers as shared information without a single invariant coordinate. Diagnostics separate the measurement from online stopping, verbalizer-free belief, and causal answer control; exact steering shows local sensitivity of $δ$ but not reliable generation control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。