让生成文本自带计算过程证据,可验证模型内部状态。
Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text

- 在生成文本中嵌入可检测的因果状态信号。
- 两种模型在128组测试中均准确传递内部状态信息。
- 适合关注AI可解释性与可信推理的研究者。
语言模型的输出本身无法提供其内部计算过程的可验证证据。本文研究计算溯源:生成文本能否携带可检测的、关于导致结果的因果相关内部状态的证据。我们在两种受控架构中验证这一设想:模块化前馈神经网络和基于Transformer的模型。两者均在相同算术任务上训练,强制通过两个离散中间状态,不同内部路径可产生相同答案。我们故意切换路径,认证实际使用的状态,并让该已验证状态决定生成文本中的细微统计模式,后续可被检测。前馈与Transformer系统在公开及独立封闭的端到端评估中,均成功通过全部128组匹配对,检测器能恢复与认证状态相关的信号。所需的因果计算也在五个独立训练的前馈模型和三个独立训练的Transformer中重现。在仅使用答案的Transformer实验中,线性探针未能恢复自然学习的中间状态。结果表明,在答案不变的情况下,关于已验证因果状态的信息仍可保留在生成文本中。
原文摘要 · Abstract (English)
A language model's output does not by itself provide verifiable evidence about the internal computation that produced it. We study computational provenance: whether generated text can carry detectable evidence of which causally relevant internal state occurred. We test a bounded form of this idea in two controlled architectures: a modular feed-forward neural network and a transformer-based model. Both architectures are trained on the same arithmetic task with a mandatory pathway through two discrete intermediate states, allowing different internal paths to produce the same answer. We deliberately switch between these paths, authenticate the state actually used, and let that verified state determine a subtle statistical pattern in the generated text that can later be detected. The feed-forward and transformer systems each passed all 128 matched pairs in both their public and separately sealed protected end-to-end evaluations, with the detector recovering the signal associated with the authenticated internal state. The required causal computation also reproduced across five independently trained feed-forward models and three independently trained transformers. In a separate answer-only transformer experiment, our linear probes did not recover a naturally learned intermediate state. These results provide a controlled proof of concept that information about a verified, causally relevant internal state can be preserved in generated text even when the answer is unchanged.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。