arXiv:2411.13612cs.CRcs.LG2024-11被引 7

针对低嵌入率短时语音流,提出双视角检测框架提升隐蔽通信识别准确率。

Efficient Streaming Voice Steganalysis in Challenging Detection Scenarios

  • 通过随机混淆原始特征,增强难检样本的隐蔽特征可学习性。
  • 在10%嵌入率、0.1秒短时条件下,检测准确率显著优于现有方法。
  • 适合实时语音安全监测场景,尤其适用于资源受限环境。

近年来,基于网络流媒体的信息隐藏技术日益增多,聚焦于如何将秘密信息高效嵌入实时传输的网络媒体信号中以实现隐蔽通信。此类技术的滥用可能带来恶意代码、指令和病毒传播等重大安全风险。当前针对网络语音流的隐写分析方法面临两大挑战:在低嵌入率(如低至10%)和短传输时长(如仅0.1秒)条件下的高效检测。由于此时检测模型难以获取足够丰富的样本特征,导致有效隐写分析困难。为此,本文提出双视角VoIP隐写分析框架(DVSF)。该框架首先对VoIP流段中的原始隐写特征进行随机混淆,使难检样本的隐写特征更明显,便于模型学习;然后捕捉与隐写相关的细粒度局部特征,并融合全局VoIP特征。特别构建的VoIP段三元组进一步调节模型内部特征距离。大量实验表明,该方法显著提升了在挑战性检测场景下流式语音隐写分析的准确率,超越现有最先进方法,并具备优异的近实时性能。

原文摘要 · Abstract (English)

In recent years, there has been an increasing number of information hiding techniques based on network streaming media, focusing on how to covertly and efficiently embed secret information into real-time transmitted network media signals to achieve concealed communication. The misuse of these techniques can lead to significant security risks, such as the spread of malicious code, commands, and viruses. Current steganalysis methods for network voice streams face two major challenges: efficient detection under low embedding rates and short duration conditions. These challenges arise because, with low embedding rates (e.g., as low as 10%) and short transmission durations (e.g., only 0.1 second), detection models struggle to acquire sufficiently rich sample features, making effective steganalysis difficult. To address these challenges, this paper introduces a Dual-View VoIP Steganalysis Framework (DVSF). The framework first randomly obfuscates parts of the native steganographic descriptors in VoIP stream segments, making the steganographic features of hard-to-detect samples more pronounced and easier to learn. It then captures fine-grained local features related to steganography, building on the global features of VoIP. Specially constructed VoIP segment triplets further adjust the feature distances within the model. Ultimately, this method effectively address the detection difficulty in VoIP. Extensive experiments demonstrate that our method significantly improves the accuracy of streaming voice steganalysis in these challenging detection scenarios, surpassing existing state-of-the-art methods and offering superior near-real-time performance.

隐写分析语音安全实时检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。