arXiv:2505.22266cs.SDcs.MM2025-05

用对抗扰动嵌入秘密信息,实现高效高保真音频隐写。

FGAS: Fixed Decoder Network-Based Audio Steganography with Adversarial Perturbation Generation

  • 固定解码器+对抗扰动生成,提升隐写鲁棒性。
  • 平均保真度提升超10 dB,抗检测能力更强。
  • 适合对隐蔽性与效率要求高的场景使用。

人工智能生成内容(AIGC)的快速发展使得高质量生成音频在互联网上广泛传播,推动了音频隐写技术的进步。当前主流音频隐写方法基于编码器-解码器架构,虽能保证一定感知质量,但普遍存在计算开销大、实现时间长及抗隐写分析性能差的问题。为此,我们提出一种基于固定解码器的音频隐写方法——FGAS(Fixed Decoder Network-Based Audio Steganography with Adversarial Perturbation Generation)。秘密信息以对抗扰动形式嵌入原始音频生成隐写音频,接收方仅需共享固定解码器的结构与密钥即可准确提取信息。我们设计了可选鲁棒性的音频对抗扰动生成策略(A2PG),并构建轻量级固定解码器。该设计确保秘密信息可靠提取,同时优化对抗扰动使隐写音频在感知和统计特性上更接近原始音频,从而提升抗检测能力。实验表明,FGAS显著提升了隐写音频质量,平均峰值信噪比(PSNR)相比现有最优方法提升超过10 dB;同时具备强健的抗常见音频处理攻击能力。此外,在不同相对载荷下均表现出优越的抗隐写分析性能:高容量嵌入时,分类错误率约高出2%,表明其隐蔽性优于当前SOTA方法。

原文摘要 · Abstract (English)

The rapid development of Artificial Intelligence Generated Content (AIGC) has made high-fidelity generated audio widely available across the Internet, driving the advancement of audio steganography. Benefiting from advances in deep learning, current audio steganography schemes are mainly based on encoder-decoder network architectures. While these methods guarantee a certain level of perceptual quality for stego audio, they typically face high computational cost and long implementation time, as well as poor anti-steganalysis performance. To address the aforementioned issues, we pioneer a Fixed Decoder Network-Based Audio Steganography with Adversarial Perturbation Generation (FGAS). Adversarial perturbations carrying a secret message are embedded into the cover audio to generate stego audio. The receiver only needs to share the structure and key of the fixed decoder network to accurately extract the secret message from the stego audio. In FGAS, we propose an Audio Adversarial Perturbation Generation (A2PG) strategy with an optional robust extension and design a lightweight fixed decoder. The fixed decoder guarantees reliable extraction of the hidden message, while adversarial perturbations are optimized to keep the stego audio perceptually and statistically close to the cover audio, thereby improving anti-steganalysis performance. The experimental results show that FGAS significantly improves stego audio quality, achieving an average PSNR gain of over 10 dB compared to SOTA methods. Furthermore, FGAS demonstrates strong robustness against common audio processing attacks. Moreover, FGAS exhibits superior anti-steganalysis performance across different relative payloads; under high-capacity embedding, it achieves a classification error rate about 2% higher, indicating stronger anti-steganalysis performance than current SOTA methods.

音频隐写对抗扰动固定解码器AIGC安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。