arXiv:2510.23035cs.CRcs.AI2025-10被引 3

用熵驱动的词元排序编码,实现高容量隐写通信

A high-capacity linguistic steganography based on entropy-driven rank-token mapping

  • 根据词元概率排名映射密文,动态调整采样策略
  • 隐蔽文本容量达主流方法3倍,处理速度提升50%以上
  • 适合对安全性和效率要求高的隐写应用

语言隐写技术通过将秘密信息嵌入无害文本实现隐蔽通信,但现有方法在容量和安全性上存在瓶颈。传统修改类方法引入可检测异常,检索类策略嵌入容量低,现代生成式隐写虽利用语言模型生成自然隐写文本,但受限于词元预测熵值不足,进一步制约容量。为此,我们提出熵驱动框架RTMStega,结合基于排名的自适应编码与上下文感知解压缩及归一化熵机制。通过将密文映射至词元概率排名,并依据上下文感知熵动态调整采样,实现了容量与不可察觉性的平衡。跨多种数据集和模型的实验表明,RTMStega将主流生成式隐写的容量提升3倍,处理时间减少超过50%,同时保持高质量文本输出,为安全高效的隐蔽通信提供了可信方案。

原文摘要 · Abstract (English)

Linguistic steganography enables covert communication through embedding secret messages into innocuous texts; however, current methods face critical limitations in payload capacity and security. Traditional modification-based methods introduce detectable anomalies, while retrieval-based strategies suffer from low embedding capacity. Modern generative steganography leverages language models to generate natural stego text but struggles with limited entropy in token predictions, further constraining capacity. To address these issues, we propose an entropy-driven framework called RTMStega that integrates rank-based adaptive coding and context-aware decompression with normalized entropy. By mapping secret messages to token probability ranks and dynamically adjusting sampling via context-aware entropy-based adjustments, RTMStega achieves a balance between payload capacity and imperceptibility. Experiments across diverse datasets and models demonstrate that RTMStega triples the payload capacity of mainstream generative steganography, reduces processing time by over 50%, and maintains high text quality, offering a trustworthy solution for secure and efficient covert communication.

隐写术语言模型信息隐藏生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。