将多种数据隐秘嵌入文字像素,视觉上几乎无痕。
Raster Domain Text Steganography: A Unified Framework for Multimodal Secure Embedding
- 在字体光栅化后修改像素,用极小改动编码信息
- 通过统计像素数量恢复数据,支持图文音视频多模态嵌入
- 方法轻量且稳定,适合普通文本做隐写载体
本文提出一种统一的光栅域隐写框架——字形扰动基数(GPC)框架,可直接将文本、图像、音频和视频等异构数据嵌入渲染后文字字形的像素空间。该方法在字体光栅化后操作,仅修改确定性渲染流程生成的位图,每个字形作为隐蔽编码单元,通过最小扰动的内部墨点数量表达载荷值。这些微弱的强度变化在视觉上不可察觉,但能形成稳定可解码的信号。框架通过归一化图像强度、音频特征和视频帧值为有界整数序列,实现对多模态输入的扩展。解码时重新光栅化原始文本,减去标准字形位图,再通过像素计数分析恢复载荷。该方法计算开销低,基于确定性光栅行为,使普通文本成为多模态数据的视觉隐形载体。
原文摘要 · Abstract (English)
This work introduces a unified raster domain steganographic framework, termed as the Glyph Perturbation Cardinality (GPC) framework, capable of embedding heterogeneous data such as text, images, audio, and video directly into the pixel space of rendered textual glyphs. Unlike linguistic or structural text based steganography, the proposed method operates exclusively after font rasterization, modifying only the bitmap produced by a deterministic text rendering pipeline. Each glyph functions as a covert encoding unit, where a payload value is expressed through the cardinality of minimally perturbed interior ink pixels. These minimal intensity increments remain visually imperceptible while forming a stable and decodable signal. The framework is demonstrated for text to text embedding and generalized to multimodal inputs by normalizing image intensities, audio derived scalar features, and video frame values into bounded integer sequences distributed across glyphs. Decoding is achieved by re-rasterizing the cover text, subtracting canonical glyph rasters, and recovering payload values via pixel count analysis. The approach is computationally lightweight, and grounded in deterministic raster behavior, enabling ordinary text to serve as a visually covert medium for multimodal data embedding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。