构建首个系统性图像隐写攻防评估框架,揭示隐写攻击与检测的不对称风险。
Stego Battlefield: Evaluating Image Steganography Attacks and Steganalysis Defenses

- 设计四类核心任务,覆盖隐写攻击、检测防御、效率与迁移性评估。
- 发现攻击在不同图像分布间迁移性强,而检测模型难以适应新分布。
- 实测社交平台隐写内容在压缩后仍可存活,威胁真实存在。
图像隐写广泛用于保护用户隐私和实现隐蔽通信,但也可能被恶意利用作为绕过内容审核、传播有害语义甚至向大模型注入危险指令的隐蔽通道,构成持续演进的实际安全风险。为填补缺乏统一系统评估框架的空白,我们提出SADBench,一个系统性基准,用于评估攻击者通过隐写注入有害秘密的能力,以及防御者通过隐写分析识别此类威胁的能力。关键的是,SADBench包含4个核心任务:隐写攻击能力评估、隐写分析防御能力评估、效率评估和迁移性评估。它在多种载体分布下,对图像载荷和文本载荷隐写进行评估,使用有害视觉语义和有毒指令模拟恶意攻击。在广泛的攻击与检测方法上,SADBench揭示:(i) INN与自编码器基方法相比其他架构具有更优稳定性;(ii) 域内检测近乎完美且比生成成本更低;(iii) 存在显著的迁移不对称性——攻击能稳健泛化至新分布,而检测器难以适应;(iv) 真实世界威胁依然存在,社交平台上载荷即使经轻微压缩仍可存活,或通过模拟训练有效适应强压缩。总体而言,SADBench建立了系统化、可复现、可扩展的评估框架,量化风险,推动可度量、以安全为导向的隐写防御发展。
原文摘要 · Abstract (English)
Image steganography is widely used to protect user privacy and enable covert communication. However, it can also be abused by the adversary as a covert channel to bypass content moderation, disseminate harmful semantics, and even hide malicious instructions in images to elicit dangerous outputs from large models, posing a practical security risk that continues to evolve. To address the lack of a unified and systematic evaluation framework, we propose SADBench, a systematic benchmark that assesses the adversary's ability to inject harmful secrets via steganography and the defender's ability to detect such threats through steganalysis. Crucially, SADBench comprises $4$ core tasks, namely steganography attack capability evaluation, steganalysis defense capability evaluation, efficiency evaluation, and transferability evaluation. It evaluates both image-payload and text-payload steganography across diverse cover distributions, utilizing harmful visual semantics and toxic instructions to simulate malicious attacks. Across a broad set of attacks and detectors, SADBench reveals that (i) INN and autoencoder-based methods demonstrate superior stability compared to other architectures, (ii) in-domain detection is near-perfect and cheaper than generation, (iii) a critical asymmetry exists in transferability where attacks robustly generalize to new distributions while detectors fail to adapt, and (iv) real-world threats persist on social media, where payloads either survive minimal compression or effectively adapt to aggressive compression via simulated training. Overall, SADBench establishes a systematic, reproducible, and extensible framework to quantify risks, paving the way for measurable and security-driven advancements in steganography defense.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。