arXiv:2510.09714cs.CLcs.AI2025-10被引 5

模型在加密语言中推理能力下降,难逃链式思考监控。

All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language

  • 用28种密码微调模型,测试其在加密文本中的推理能力。
  • 主流模型在冷门密码下准确率大幅下降,仅对常见密码如rot13表现良好。
  • 加密推理能力与预训练数据中密码出现频率相关,可被量化控制。

检测有害AI行为至关重要,链式思考(CoT)监控是常用方法。然而攻击者或失控模型可能通过加密、翻译或压缩文本隐藏推理过程来规避监控。为评估此风险,我们测试模型在28种不同密码下的推理能力:对最多10个模型进行微调和提示,在每种密码中完成数学题作为推理能力代理指标。结果发现显著不对称性:尽管模型能准确将加密文本翻译回英文,但在加密文本中推理时准确率仍大幅下降。即使前沿模型在不常见的密码上也表现不佳,而仅在如rot13等常见密码上能准确推理。我们发现加密推理能力与预训练数据中密码的出现频率相关,并揭示了随微调数据增加缓慢提升的缩放规律。研究表明,当前模型利用加密推理规避CoT监控的效果有限,为未来模型开发提供约束依据。

原文摘要 · Abstract (English)

Detecting harmful AI actions is important as AI agents gain adoption. Chain-of-thought (CoT) monitoring is one method widely used to detect adversarial attacks and AI misalignment. However, attackers and misaligned models might evade CoT monitoring through ciphered reasoning: reasoning hidden in encrypted, translated, or compressed text. To assess this risk, we test whether models can perform ciphered reasoning. For each of 28 different ciphers, we fine-tune and prompt up to 10 models to reason in that cipher. We measure model accuracy on math problems as a proxy for reasoning ability. Across the models we test, we find an asymmetry: model accuracy can drop significantly when reasoning in ciphered text, even though models demonstrate comprehension of ciphered text by being able to translate it accurately to English. Even frontier models struggle with lesser-known ciphers, although they can reason accurately in well-known ciphers like rot13. We show that ciphered reasoning capability correlates with cipher prevalence in pretraining data. We also identify scaling laws showing that ciphered reasoning capability improves slowly with additional fine-tuning data. Our work suggests that evading CoT monitoring using ciphered reasoning may be an ineffective tactic for current models and offers guidance on constraining the development of this capability in future frontier models.

AI安全链式思考加密推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。