arXiv:2606.09411cs.CRcs.IT2026-06

提出可绕过现有检测的隐写攻击,并通过数据干预恢复检测能力。

Now You (Still) See Me: Detecting Evasive Steganographic Payloads in LLMs

论文配图:Now You (Still) See Me: Detecting Evasive Steganographic Payloads in LLMs
图 1 · 摘自论文原文
  • 用非线性探针增强检测,但攻击仍能绕过。
  • 隐写模型保留58%-79%秘密还原率,仅损失1%-8%性能。
  • 重构数据分布可让隐蔽信息重新暴露,适合安全研究者。

大型语言模型可通过微调将提示中的秘密编码为流畅且看似无害的输出,构成难以通过输出层面分析检测的隐写泄露风险。已有研究采用线性探针从内部激活中恢复秘密,实现机制检测。我们发现该防御可被系统性规避,但通过针对性的数据级干预可恢复检测能力。首先,我们将检测设置扩展至非线性MLP探针;随后,在五个基座模型(Qwen3-8B、Llama-3.1-8B、Ministral-8B、Qwen3-14B、Phi-4-14B)上对抗性微调隐写木马,结果模型在六项基准测试中平均性能下降1%-8%,但保留58%-79%的精确秘密恢复率,成功规避了岭回归与预留的MLP探针。我们进一步从信息论角度分析该规避现象:成功规避保持秘密可恢复性,同时降低内容对齐表示中秘密的低阶可提取性,迫使载荷与残差自由度产生协同作用。据此提出重上下文数据集以限制残差自由度,在该分布下,所有五种规避型木马的岭回归与MLP检测能力均得以恢复。总体表明,基于激活的隐写检测易受自适应攻击影响,但理论指导下的评估分布可揭示隐藏载荷。

原文摘要 · Abstract (English)

Large language models can be fine-tuned to encode prompt-borne secrets into fluent, seemingly benign outputs. This creates a steganographic exfiltration risk that is difficult to detect with output-level steganalysis. Recent work proposes mechanistic detection using linear probes that recover the secret from internal activations. We show that this defense can be systematically evaded, but that detectability can be recovered through a targeted data-level intervention. First, we extend the detection setup to include a non-linear MLP probe. We then adversarially fine-tune steganographic trojans across five base models: Qwen3-8B, Llama-3.1-8B, Ministral-8B, Qwen3-14B, and Phi-4-14B. The resulting models retain $58$--$79\%$ exact-match secret recovery while evading both ridge and held-out MLP probes, with $1$--$8\%$ average capability degradation across six benchmarks. We then give an information-theoretic characterization of this evasion. Successful evasion preserves recoverability while reducing low-order extractability of the secret from the content-aligned representation, forcing the payload into synergistic interaction with residual degrees of freedom. This motivates a recontextualization dataset that restricts these residual degrees of freedom. On this distribution, both ridge and MLP detectability are restored across all five evasive trojans. Overall, our findings show that activation-based steganography detection is vulnerable to adaptive evasion, but also that theory-guided evaluation distributions can expose otherwise hidden payloads.

隐写攻击模型安全检测防御大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。