构建新数据集揭示神经语音编辑生成的局部伪造音频难被现有模型检测
PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
- 用先进神经编辑技术生成局部篡改语音,构建新数据集PartialEdit
- 现有检测模型在新数据上准确率骤降,无法识别神经编码器生成的伪造痕迹
- 适合语音安全、深度伪造检测方向的研究者关注
神经语音编辑可对语音片段进行精准局部修改,同时保持其余部分不变,这一技术虽具实用性,但也带来新型深度伪造风险。为推动对此类局部伪造语音的检测研究,本文提出PartialEdit数据集,其通过先进神经编辑技术构建。我们在该数据集上探索了检测与定位任务,实验表明,基于现有PartialSpoof数据集训练的模型无法有效识别由神经语音编辑模型生成的局部篡改语音。由于近期语音编辑模型普遍采用神经音频编码器,本文还分析了模型学习到的检测特征与潜在伪造痕迹。更多关于PartialEdit数据集及音频样本的信息请访问项目主页:https://yzyouzhang.com/PartialEdit/index.html。
原文摘要 · Abstract (English)
Neural speech editing enables seamless partial edits to speech utterances, allowing modifications to selected content while preserving the rest of the audio unchanged. This useful technique, however, also poses new risks of deepfakes. To encourage research on detecting such partially edited deepfake speech, we introduce PartialEdit, a deepfake speech dataset curated using advanced neural editing techniques. We explore both detection and localization tasks on PartialEdit. Our experiments reveal that models trained on the existing PartialSpoof dataset fail to detect partially edited speech generated by neural speech editing models. As recent speech editing models almost all involve neural audio codecs, we also provide insights into the artifacts the model learned on detecting these deepfakes. Further information about the PartialEdit dataset and audio samples can be found on the project page: https://yzyouzhang.com/PartialEdit/index.html.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。