不修改原文就能藏密信,还能100%准确提取。
A Content-Preserving Secure Linguistic Steganography
- 用可调概率分布替换原文词序,隐藏信息不改内容。
- 实验显示密文提取成功率100%,安全性优于现有方法。
- 适合对内容完整性要求高的安全通信场景。
现有语言隐写方法多通过内容变换隐藏秘密信息,但常导致正常文本与隐写文本间出现细微却可疑的差异,带来真实应用中的安全风险。为此,我们提出一种内容无损的语言隐写范式,实现完全安全的隐蔽通信,且不修改原始载体文本。基于此范式,我们设计了CLstega(Content-preserving Linguistic steganography),通过可控的概率分布转换嵌入秘密信息。CLstega首先采用增强掩码策略定位并标记嵌入位置,利用MLM(掩码语言模型)预测的概率分布易于调整以实现变换;随后设计动态分布隐写编码策略,从原始概率分布推导目标分布来编码秘密信息。为实现该变换,CLstega精心选择嵌入位置的目标词作为标签,构建掩码句子数据集,用于微调原始MLM,生成能直接从载体文本中提取秘密信息的目标MLM。该方法确保秘密信息的完美安全性,同时完全保留原始覆盖文本的完整性。实验结果表明,CLstega可实现100%的提取成功率,在安全性方面优于现有方法,有效平衡了嵌入容量与安全性。
原文摘要 · Abstract (English)
Existing linguistic steganography methods primarily rely on content transformations to conceal secret messages. However, they often cause subtle yet looking-innocent deviations between normal and stego texts, posing potential security risks in real-world applications. To address this challenge, we propose a content-preserving linguistic steganography paradigm for perfectly secure covert communication without modifying the cover text. Based on this paradigm, we introduce CLstega (\textit{C}ontent-preserving \textit{L}inguistic \textit{stega}nography), a novel method that embeds secret messages through controllable distribution transformation. CLstega first applies an augmented masking strategy to locate and mask embedding positions, where MLM(masked language model)-predicted probability distributions are easily adjustable for transformation. Subsequently, a dynamic distribution steganographic coding strategy is designed to encode secret messages by deriving target distributions from the original probability distributions. To achieve this transformation, CLstega elaborately selects target words for embedding positions as labels to construct a masked sentence dataset, which is used to fine-tune the original MLM, producing a target MLM capable of directly extracting secret messages from the cover text. This approach ensures perfect security of secret messages while fully preserving the integrity of the original cover text. Experimental results show that CLstega can achieve a 100\% extraction success rate, and outperforms existing methods in security, effectively balancing embedding capacity and security.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。