应对语音深度伪造威胁,提出三类防护技术。
Speaker Privacy and Security in the Big Data Era: Protection and Defense against Deepfake
- 通过语音匿名化保护声纹特征不被提取
- 结合检测与水印技术识别并防范深度伪造语音
- 适合关注语音安全与隐私的科研及产业人士
在大数据时代,个性化语音生成技术取得了显著进展,利用说话人特征(如声音和语调)生成深度伪造语音。这加剧了全球范围内的语音滥用安全风险,带来巨大社会成本。为应对深度伪造语音威胁,研究聚焦于语音属性保护与防御两大方向:语音匿名化技术可防止声纹特征被提取用于伪造;深度伪造检测与水印技术则用于识别和防范伪造语音的滥用。本文简要综述了这三项技术的方法、进展与挑战,完整版本将不久后发布。
原文摘要 · Abstract (English)
In the era of big data, remarkable advancements have been achieved in personalized speech generation techniques that utilize speaker attributes, including voice and speaking style, to generate deepfake speech. This has also amplified global security risks from deepfake speech misuse, resulting in considerable societal costs worldwide. To address the security threats posed by deepfake speech, techniques have been developed focusing on both the protection of voice attributes and the defense against deepfake speech. Among them, the voice anonymization technique has been developed to protect voice attributes from extraction for deepfake generation, while deepfake detection and watermarking have been utilized to defend against the misuse of deepfake speech. This paper provides a short and concise overview of the three techniques, describing the methodologies, advancements, and challenges. A comprehensive version, offering additional discussions, will be published in the near future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。