用决策理论定义隐写术,可检测大模型隐藏指令
A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring
- 以可用信息差异为视角,构建隐写检测新框架
- 提出'隐写差距'指标,量化模型隐藏信息的能力
- 适合关注大模型安全与监控的研究者
大型语言模型开始展现隐写能力,可能使对齐不良的模型规避监管机制。然而,现有隐写检测方法依赖已知的非隐写信号分布,而大模型隐写推理中这一分布无法获取,导致传统方法失效。本文提出一种决策理论视角下的隐写形式化方法,核心洞察是:能解码隐藏内容的代理与不能解码的代理之间存在可用信息不对称,这种潜在差异可通过可观测行为推断。为此,我们引入广义$$\mathcal{V}$-信息——一种衡量输入中可用信息量的功利性框架,并基于此定义了'隐写差距',即比较可解码与不可解码代理在下游任务中的效用差异,从而量化隐写程度。我们在实验中验证了该形式化的有效性,证明其可用于检测、量化和缓解大模型中的隐写推理行为。
原文摘要 · Abstract (English)
Large language models are beginning to show steganographic capabilities. Such capabilities could allow misaligned models to evade oversight mechanisms. Yet principled methods to detect and quantify such behaviours are lacking. Classical definitions of steganography, and detection methods based on them, require a known reference distribution of non-steganographic signals. For the case of steganographic reasoning in LLMs, knowing such a reference distribution is not feasible; this renders these approaches inapplicable. We propose an alternative, \textbf{decision-theoretic view of steganography}. Our central insight is that steganography creates an asymmetry in usable information between agents who can and cannot decode the hidden content (present within a steganographic signal), and this otherwise latent asymmetry can be inferred from the agents' observable actions. To formalise this perspective, we introduce generalised $\mathcal{V}$-information: a utilitarian framework for measuring the amount of usable information within some input. We use this to define the \textbf{steganographic gap} -- a measure that quantifies steganography by comparing the downstream utility of the steganographic signal to agents that can and cannot decode the hidden content. We empirically validate our formalism, and show that it can be used to detect, quantify, and mitigate steganographic reasoning in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。