用可逆神经网络实现高质量语音隐写,恢复音质更好且抗干扰。
WavInWav: Time-domain Speech Hiding via Invertible Neural Network
- 基于可逆神经网络建立语音隐写三者直接关联,提升信息可逆性。
- 在时域引入时频损失,使秘密语音恢复质量显著优于以往方法。
- 支持加密保护,适合对音质和安全有要求的通信场景。
数据隐藏对数字媒体中的安全通信至关重要,深度神经网络为有效嵌入秘密信息提供了新方法。然而,以往音频隐写方法因难以建模时频关系,导致秘密语音恢复质量不佳。本文探索该局限,提出一种基于流模型的可逆神经网络方法,建立伪音频、原始音频与秘密音频之间的直接联系,增强嵌入与提取的可逆性。为克服时频变换导致的恢复质量下降问题,我们在时域信号上引入时频损失,既保留时频约束优势,又提升消息恢复的可逆性,这对实际应用至关重要。此外,还加入加密技术保护隐藏数据。在VCTK和LibriSpeech数据集上的实验表明,本方法在主观与客观指标上均优于现有方法,并对各类噪声具有鲁棒性,适用于特定安全通信场景。
原文摘要 · Abstract (English)
Data hiding is essential for secure communication across digital media, and recent advances in Deep Neural Networks (DNNs) provide enhanced methods for embedding secret information effectively. However, previous audio hiding methods often result in unsatisfactory quality when recovering secret audio, due to their inherent limitations in the modeling of time-frequency relationships. In this paper, we explore these limitations and introduce a new DNN-based approach. We use a flow-based invertible neural network to establish a direct link between stego audio, cover audio, and secret audio, enhancing the reversibility of embedding and extracting messages. To address common issues from time-frequency transformations that degrade secret audio quality during recovery, we implement a time-frequency loss on the time-domain signal. This approach not only retains the benefits of time-frequency constraints but also enhances the reversibility of message recovery, which is vital for practical applications. We also add an encryption technique to protect the hidden data from unauthorized access. Experimental results on the VCTK and LibriSpeech datasets demonstrate that our method outperforms previous approaches in terms of subjective and objective metrics and exhibits robustness to various types of noise, suggesting its utility in targeted secure communication scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。