无需目标水印算法即可移除音频水印,且泛化能力强。
HarmonicAttack: An Adaptive Cross-Domain Audio Watermark Removal
- 通过少量原始与加水印音频训练通用去水印模型,无需知晓水印算法。
- 在多个数据集和水印方案上均实现超90%的识别成功率,远超现有方法。
- 适用于评估水印鲁棒性,尤其适合研究安全防御的团队。
高质量AI生成音频的普及带来了虚假信息传播和语音克隆欺诈等安全挑战。音频水印是防范滥用的关键手段,但攻击者可能尝试移除水印,因此研究有效移除技术对客观评估水印鲁棒性至关重要。以往方法通常假设可访问目标水印检测器,这一假设常不现实,导致对现有水印方案过于乐观。本文提出HarmonicAttack,一种新型音频水印移除方法,无需目标水印算法,仅需少量原始与加水印样本即可训练通用模型。我们发现训练样本无需与目标样本同分布,其攻击在分布外样本上仍表现稳定,性能下降极小。相比现有方法,HarmonicAttack在AudioSeal、WavMark、SilentCipher和AudioMarkNet等先进水印方案中均更有效,同时保持高听觉质量。尽管在LibriSpeech数据集上针对AudioSeal训练,其仍能跨数据集(如VCTK)和水印方案泛化:在VCTK上对AudioMarkNet实现92%的自动语音识别成功率(ASR),远超最佳基线38%;在FMA上对所有水印达100% ASR,而最佳基线仅2%(AudioSeal)和44%(WavMark)。
原文摘要 · Abstract (English)
The availability of high-quality, AI-generated audio raises security challenges such as misinformation campaigns and voice-cloning fraud. A key defense against the misuse of AI-generated audio is by watermarking it, so that it can be easily distinguished from genuine audio. Those seeking to misuse AI-generated audio may attempt to remove audio watermarks, so studying effective watermark removal techniques is critical to objectively evaluate the robustness of audio watermarks. Previous watermark removal schemes typically assume access to the target watermark detector during the removal process. This assumption is often impractical, which may lead to a false sense of confidence in current watermark schemes. We introduce HarmonicAttack, a novel audio watermark removal method that requires no access to the target watermark algorithm. It only needs a number of original and watermarked samples to train a general model capable of removing watermarks from audio samples. We also find that training samples do not need to share the same distribution as target samples, as our attack generalizes to out-of-distribution samples with minimal degradation. Compared with existing watermark removal attacks, HarmonicAttack is more effective at removing watermarks from state-of-the-art schemes, including AudioSeal, WavMark, SilentCipher, and AudioMarkNet, while maintaining high perceptual quality. Although HarmonicAttack is trained on the LibriSpeech dataset against AudioSeal, it generalizes across unseen datasets and watermarking schemes. For instance, on VCTK, HarmonicAttack achieves a 92% ASR against AudioMarkNet, substantially outperforming the best baseline at 38%. On FMA, HarmonicAttack reaches 100% ASR against all watermarks, whereas the best baseline achieves only 2% against AudioSeal and 44% against WavMark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。