arXiv:2410.04328cs.ITcs.AI2024-10Conference of the …被引 8

用优化分布让大模型生成更隐蔽的隐写文本,减少用词量且保持自然。

OD-Stega: LLM-Based Relatively Secure Steganography via Optimized Distributions

  • 通过优化下一个词的概率分布,实现高效隐写。
  • 在保持文本自然性的前提下,用最少的词嵌入秘密信息。
  • 解决分词不匹配等问题,兼容现有隐写技术。

我们研究无载体隐写,利用大语言模型(LLM)生成隐写文本,并结合算术编码。高效的方法应在尽可能少的语言标记中嵌入秘密比特,同时使隐写文本尽可能自然。我们证明该问题等价于在限制新分布与原始分布之间散度的前提下,最大化下一个词生成的概率分布熵。在KL散度或总变差约束下,给出了闭式解。还解决了几个实际问题:1)通过简单提示选择方法解决了常见的分词不匹配问题;2)考虑了优化分布与词汇截断技术的结合;3)将所提方法与现有(可能非算术编码基础的)技术(如Discop方法)融合。

原文摘要 · Abstract (English)

We consider coverless steganography where a Large Language Model (LLM) is used to generate stego-texts in combination with arithmetic coding. An efficient method should embed secret bits in as few language tokens as possible while keeping the stego-text as natural as possible. We show that this problem is equivalent to maximizing the entropy of a replacement probability distribution of the next token generation, subject to a constraint on the divergence between the new distribution and the original one produced by the LLM. A closed-form solution is provided under either the KL divergence or the total variation constraint. Several important practical issues are also tackled: 1) An often-overlooked tokenization mismatch issue is resolved with a simple prompt selection approach, 2) The combination of the optimized distribution and the vocabulary truncation technique is considered, and 3) The incorporation of the proposed approach with existing (potentially non arithmetic coding based) techniques, e.g., the Discop technique.

隐写大模型语言模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。