arXiv:2603.21567cs.LG2026-03被引 1

揭示大模型隐写术的信息论代价,并提出用困惑度比检测隐藏信息。

Kolmogorov Complexity Bounds for LLM Steganography and a Perplexity-Based Detection Proxy

  • 基于柯尔莫哥洛夫复杂度证明:隐写文本复杂度必高于原文加密内容
  • 实验证明隐写文本困惑度显著升高,差异具有统计显著性(p < 10⁻⁶)
  • 适合关注AI安全、隐写检测与模型对齐的研究者阅读

大语言模型可重写文本以嵌入隐藏信息,同时保持表面语义不变,这为协作式AI系统开辟了隐蔽通信通道,并对对齐监控构成挑战。本文研究此类嵌入的信息论成本。主要结论表明:任何在保持原文语义负载 $M_1$ 的前提下,将载荷 $P$ 编码至隐写文本 $M_2$ 的隐写方案,必须满足 $K(M_2) \ geq K(M_1) + K(P) - O(\log n)$,其中 $K$ 表示柯尔莫哥洛夫复杂度,$n$ 为总消息长度。由此推论:只要载荷非平凡,隐写文本的复杂度必然严格上升,无论编码器如何分布信号。由于柯尔莫哥洛夫复杂度不可计算,我们探讨是否可用实用代理检测这一预测的增长。借鉴无损压缩与柯尔莫哥洛夫复杂度的经典对应关系,我们主张语言模型困惑度在概率框架中扮演类似角色,并提出‘双筒望远镜’困惑度比作为代理指标。针对基于颜色的LLM隐写方案的初步实验支持理论预测:300个样本的配对t检验显示 $t = 5.11$,$p < 10^{-6}$。

原文摘要 · Abstract (English)

Large language models can rewrite text to embed hidden payloads while preserving surface-level meaning, a capability that opens covert channels between cooperating AI systems and poses challenges for alignment monitoring. We study the information-theoretic cost of such embedding. Our main result is that any steganographic scheme that preserves the semantic load of a covertext~$M_1$ while encoding a payload~$P$ into a stegotext~$M_2$ must satisfy $K(M_2) \geq K(M_1) + K(P) - O(\log n)$, where $K$ denotes Kolmogorov complexity and $n$ is the combined message length. A corollary is that any non-trivial payload forces a strict complexity increase in the stegotext, regardless of how cleverly the encoder distributes the signal. Because Kolmogorov complexity is uncomputable, we ask whether practical proxies can detect this predicted increase. Drawing on the classical correspondence between lossless compression and Kolmogorov complexity, we argue that language-model perplexity occupies an analogous role in the probabilistic regime and propose the Binoculars perplexity-ratio score as one such proxy. Preliminary experiments with a color-based LLM steganographic scheme support the theoretical prediction: a paired $t$-test over 300 samples yields $t = 5.11$, $p < 10^{-6}$.

隐写术语言模型安全检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。