arXiv:2512.01222cs.AI2025-12被引 4

用语言模型可解释性技术破解加密推理过程,验证了当前方法的鲁棒性。

Unsupervised decoding of encoded reasoning using language model interpretability

  • 在旋转13位加密的推理中,通过内部激活值还原真实思维链。
  • 对数透镜在中间到后期层解码准确率最高,接近完整还原。
  • 无需标注数据,实现端到端自动重构推理过程,适合安全审计场景。

随着大语言模型能力增强,其可能发展出人类无法直接观察的隐藏推理过程。为检验现有可解释性技术是否能穿透此类编码推理,我们构建了一个受控测试环境:微调 DeepSeek-R1-Distill-Llama-70B 模型,在使用 ROT-13 加密进行链式思考的同时保持英文输出可读。评估机制可解释性方法,特别是对数透镜分析(logit lens),仅基于内部激活值解码模型隐藏推理过程的能力。结果显示,对数透镜能有效翻译加密推理,准确率在中间至后期层达到峰值。最终,我们提出一个完全无监督的解码流程,结合对数透镜与自动重述技术,成功从模型内部表征中重建完整的推理轨迹。这些发现表明,当前机制可解释性方法对简单形式的编码推理具备比以往认知更强的鲁棒性。本研究为评估可解释性方法应对非人可读推理格式提供了初步框架,助力维护日益强大的人工智能系统的可控性。

原文摘要 · Abstract (English)

As large language models become increasingly capable, there is growing concern that they may develop reasoning processes that are encoded or hidden from human oversight. To investigate whether current interpretability techniques can penetrate such encoded reasoning, we construct a controlled testbed by fine-tuning a reasoning model (DeepSeek-R1-Distill-Llama-70B) to perform chain-of-thought reasoning in ROT-13 encryption while maintaining intelligible English outputs. We evaluate mechanistic interpretability methods--in particular, logit lens analysis--on their ability to decode the model's hidden reasoning process using only internal activations. We show that logit lens can effectively translate encoded reasoning, with accuracy peaking in intermediate-to-late layers. Finally, we develop a fully unsupervised decoding pipeline that combines logit lens with automated paraphrasing, achieving substantial accuracy in reconstructing complete reasoning transcripts from internal model representations. These findings suggest that current mechanistic interpretability techniques may be more robust to simple forms of encoded reasoning than previously understood. Our work provides an initial framework for evaluating interpretability methods against models that reason in non-human-readable formats, contributing to the broader challenge of maintaining oversight over increasingly capable AI systems.

可解释性推理还原无监督学习模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。