用大模型将句子直接藏进图片,还能准确取出
$\mathbf{S^2LM}$: Towards Semantic Steganography via Large Language Models
- 用大模型全流程设计隐写,直接嵌入整句信息
- 在新基准IVT上实现句子级内容的准确恢复
- 适合需要隐藏语义文本的隐私与安全应用
尽管隐写技术取得显著进展,但如何在载体中嵌入具有丰富语义、句级层次的信息仍是难题。本文提出语义隐写新概念,旨在隐藏有意义的结构化内容(如句子或段落)。以图像为载体,构建句子到图像的隐写方法,实现任意句级消息的嵌入。为此提出S^2LM:基于大语言模型的语义隐写模型,全程利用LLM嵌入高层文本信息,突破传统比特级隐写的局限。此外,建立名为Invisible Text(IVT)的新基准,包含多样句级秘密文本,用于评估语义隐写方法。实验表明,S^2LM可实现超越比特级隐写的直接句子恢复。代码与数据集即将开源。
原文摘要 · Abstract (English)
Despite remarkable progress in steganography, embedding semantically rich, sentence-level information into carriers remains a challenging problem. In this work, we present a novel concept of Semantic Steganography, which aims to hide semantically meaningful and structured content, such as sentences or paragraphs, in cover media. Based on this concept, we present Sentence-to-Image Steganography as an instance that enables the hiding of arbitrary sentence-level messages within a cover image. To accomplish this feat, we propose S^2LM: Semantic Steganographic Language Model, which leverages large language models (LLMs) to embed high-level textual information into images. Unlike traditional bit-level approaches, S^2LM redesigns the entire pipeline, involving the LLM throughout the process to enable the hiding and recovery of arbitrary sentences. Furthermore, we establish a benchmark named Invisible Text (IVT), comprising a diverse set of sentence-level texts as secret messages to evaluate semantic steganography methods. Experimental results demonstrate that S^2LM effectively enables direct sentence recovery beyond bit-level steganography. The source code and IVT dataset will be released soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。