通过字符级扩散机制提升隐写文本生成质量,让隐藏信息更难被察觉。
A Character-based Diffusion Embedding Algorithm for Enhancing the Generation Quality of Generative Linguistic Steganographic Texts
- 基于字符级统计与幂律分组,提升高概率词选择频率
- 在长序列中结合XLNet模型,显著改善隐写文本流畅性
- 适合关注隐写文本质量与隐蔽性的研究人员
生成高质量隐写文本是生成式语言隐写领域的核心挑战,主要源于两方面:现有模型生成能力有限,以及嵌入算法无法有效缓解敏感信息(如语义内容或随机性)的负面影响。为确保接收方能准确提取隐藏信息,嵌入算法常需选取低概率候选词,导致高概率词减少、低概率词增多,破坏文本语义连贯性与逻辑流畅性,降低整体生成质量。本文提出一种新型嵌入算法——字符级扩散嵌入算法(CDEA)。不同于传统方法试图消除敏感信息影响,CDEA主动利用其特性,基于字符级通用统计特征与幂律分布分组方法,提升高概率候选词在候选池中的选择频率,同时降低低概率词的使用频率。此外,为保障长序列中敏感信息的有效转换,引入XLNet模型。实验表明,CDEA与XLNet结合可显著提升生成隐写文本的质量,尤其在感知不可察觉性方面表现突出。
原文摘要 · Abstract (English)
Generating high-quality steganographic text is a fundamental challenge in the field of generative linguistic steganography. This challenge arises primarily from two aspects: firstly, the capabilities of existing models in text generation are limited; secondly, embedding algorithms fail to effectively mitigate the negative impacts of sensitive information's properties, such as semantic content or randomness. Specifically, to ensure that the recipient can accurately extract hidden information, embedding algorithms often have to consider selecting candidate words with relatively low probabilities. This phenomenon leads to a decrease in the number of high-probability candidate words and an increase in low-probability candidate words, thereby compromising the semantic coherence and logical fluency of the steganographic text and diminishing the overall quality of the generated steganographic material. To address this issue, this paper proposes a novel embedding algorithm, character-based diffusion embedding algorithm (CDEA). Unlike existing embedding algorithms that strive to eliminate the impact of sensitive information's properties on the generation process, CDEA leverages sensitive information's properties. It enhances the selection frequency of high-probability candidate words in the candidate pool based on general statistical properties at the character level and grouping methods based on power-law distributions, while reducing the selection frequency of low-probability candidate words in the candidate pool. Furthermore, to ensure the effective transformation of sensitive information in long sequences, we also introduce the XLNet model. Experimental results demonstrate that the combination of CDEA and XLNet significantly improves the quality of generated steganographic text, particularly in terms of perceptual-imperceptibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。