AI语音生成与检测技术综述,揭示其应用与安全风险
A survey of AI-generated voices and their detection
- 系统梳理语音生成与检测的技术原理与最新进展
- 指出语音伪造易引发诈骗、误导等现实危害
- 适合关注AI安全、语音识别与内容可信性的研究者
人工智能模型生成高度逼真人声的能力迅速提升,广泛应用于无障碍工具、虚拟助手和创意场景,但也被用于身份冒用、欺诈和虚假信息传播。近期针对企业和政要的语音克隆诈骗事件凸显了建立强大防护机制的紧迫性。与图像和视频深度伪造不同,合成语音检测因语音的音素、语调和听觉感知复杂性而面临独特挑战。本综述全面概述了AI语音生成与检测方法,涵盖技术基础与最新前沿成果,并识别关键开放问题、基准资源与未来方向,旨在为后续研究提供实用参考。
原文摘要 · Abstract (English)
The ability of artificial intelligence (AI) models to generate highly realistic human voices has advanced rapidly. These technologies power accessibility tools, virtual assistants and creative applications, but they also enable harmful uses, including impersonation, fraud and disinformation. Recent incidents of voice cloning scams targeting businesses and political leaders underscore the urgent need for robust safeguards. Unlike image and video deepfakes, the detection of synthetic voices poses unique challenges due to the complexity of phonetics, prosody and auditory perception. This survey offers a comprehensive overview of AI voice generation and detection methods, encompassing both the technical foundations and the latest state-of-the-art advances. This study also identifies key open challenges, benchmark resources and future directions to make this survey useful for future researchers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。