用洗劫攻击增强数据,提升语音伪造检测在复杂环境下的表现
Augmentation through Laundering Attacks for Audio Spoof Detection
- 通过模拟洗劫攻击生成对抗样本,增强训练数据多样性
- 在ASVspoof 5挑战中对特定攻击(A18-A20、A26、A30)和编码条件(C08-C10)表现最差
- 适合关注实际场景下语音伪造检测鲁棒性的研究者
近期文本转语音技术的发展使语音克隆更加逼真、廉价且易获取,引发诸多滥用风险,如拜登新罕布什尔州深度伪造电话事件。已有多种检测方法被提出,但多基于相对干净的数据集训练与评估。ASVspoof 5挑战引入了一个新型众包数据库,涵盖多样声学条件、多种伪造攻击及编解码环境。本文为该挑战的提交方案,旨在研究通过洗劫攻击进行数据增强的语音伪造检测系统在ASVspoof 5数据库上的性能。结果表明,该系统在A18、A19、A20、A26和A30等攻击类型,以及C08、C09和C10编解码条件下表现最差。
原文摘要 · Abstract (English)
Recent text-to-speech (TTS) developments have made voice cloning (VC) more realistic, affordable, and easily accessible. This has given rise to many potential abuses of this technology, including Joe Biden's New Hampshire deepfake robocall. Several methodologies have been proposed to detect such clones. However, these methodologies have been trained and evaluated on relatively clean databases. Recently, ASVspoof 5 Challenge introduced a new crowd-sourced database of diverse acoustic conditions including various spoofing attacks and codec conditions. This paper is our submission to the ASVspoof 5 Challenge and aims to investigate the performance of Audio Spoof Detection, trained using data augmentation through laundering attacks, on the ASVSpoof 5 database. The results demonstrate that our system performs worst on A18, A19, A20, A26, and A30 spoofing attacks and in the codec and compression conditions of C08, C09, and C10.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。