arXiv:2410.20825cs.CRcs.AI2024-10被引 2

通过信息熵控制提升隐写文本质量与隐蔽性

ADLM -- stega: A Universal Adaptive Token Selection Algorithm for Improving Steganographic Text Quality via Information Entropy

  • 基于信息熵约束设计自适应词表筛选算法
  • 控制熵值范围可显著提升文本质量和抗检测能力
  • 适合关注隐私保护与自然语言隐写的开发者

在信息广泛共享的背景下,信息安全与隐私保护日益重要。隐写系统通过将机密信息嵌入公开载体来增强安全性;然而,现有生成式文本隐写方法难以应对候选词池的长尾分布问题,影响隐写内容的隐蔽性。本文提出一种基于信息熵约束的隐写文本质量控制理论,探索隐写文本隐蔽性与信息熵之间的关系。通过将候选词池的信息熵控制在特定范围内,优化隐写文本的隐蔽性。我们建立了信息熵的上下界,并引入自适应截断方法,在语义连贯性与词汇多样性间取得平衡。实验表明,合理控制候选词池规模和信息熵阈值能显著提升隐写文本的质量与抗检测能力,展现出在自然语言处理领域的广泛应用前景。

原文摘要 · Abstract (English)

In the context of widespread global information sharing, information security and privacy protection have become focal points. Steganographic systems enhance information security by embedding confidential information into public carriers; however, existing generative text steganography methods face challenges in handling the long-tail distribution of candidate word pools, which impacts the imperceptibility of steganographic information. This paper proposes a quality control theory for steganographic text generation based on information entropy constraints, exploring the relationship between the imperceptibility of steganographic texts and information entropy. By controlling the information entropy of the candidate word pool within a specific range, we optimize the imperceptibility of the steganographic text. We establish upper and lower bounds for information entropy and introduce an adaptive truncation method to balance semantic coherence and lexical diversity. Experimental results demonstrate that reasonably controlling the candidate pool size and information entropy thresholds significantly enhances the quality and detection resistance of steganographic texts, showcasing broad application potential in the field of natural language processing.

隐写术信息熵文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。