用音频水印技术自验证语音真伪,防篡改且抗压缩
SpeechVerifier: Robust Acoustic Fingerprint against Tampering Attacks via Watermarking
- 多尺度特征+对比学习生成抗干扰指纹
- 在无参考数据下仍能检测篡改,对压缩等操作鲁棒
- 自嵌入水印实现无需外部数据的完整性验证
随着社交媒体兴起,对重要人物公开演讲的恶意篡改严重威胁社会稳定与公众信任。现有检测方法或依赖外部参考数据,或难以同时兼顾对攻击的敏感性与对正常操作(如压缩、重采样)的鲁棒性。为此,我们提出SpeechVerifier,仅凭发布后的语音本身即可主动验证其完整性,无需外部参考。受音频指纹与水印启发,SpeechVerifier通过多尺度特征提取捕捉不同时间粒度的语音特征,并利用对比学习生成可检测多种粒度修改的指纹。该指纹对正常操作保持鲁棒,但遭遇恶意篡改时会产生显著变化。为实现自包含验证,指纹通过分段水印嵌入语音信号。发布后可直接从音频中提取指纹并与嵌入水印比对,以验证完整性。大量实验表明,该方法在检测篡改攻击方面有效,且对良性操作具有强鲁棒性。
原文摘要 · Abstract (English)
With the surge of social media, maliciously tampered public speeches, especially those from influential figures, have seriously affected social stability and public trust. Existing speech tampering detection methods remain insufficient: they either rely on external reference data or fail to be both sensitive to attacks and robust to benign operations, such as compression and resampling. To tackle these challenges, we introduce SpeechVerifer to proactively verify speech integrity using only the published speech itself, i.e., without requiring any external references. Inspired by audio fingerprinting and watermarking, SpeechVerifier can (i) effectively detect tampering attacks, (ii) be robust to benign operations and (iii) verify the integrity only based on published speeches. Briefly, SpeechVerifier utilizes multiscale feature extraction to capture speech features across different temporal resolutions. Then, it employs contrastive learning to generate fingerprints that can detect modifications at varying granularities. These fingerprints are designed to be robust to benign operations, but exhibit significant changes when malicious tampering occurs. To enable speech verification in a self-contained manner, the generated fingerprints are then embedded into the speech signal by segment-wise watermarking. Without external references, SpeechVerifier can retrieve the fingerprint from the published audio and check it with the embedded watermark to verify the integrity of the speech. Extensive experimental results demonstrate that the proposed SpeechVerifier is effective in detecting tampering attacks and robust to benign operations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。