揭示语音伪造检测中嵌入表示的可解释性,发现其保留了性别、语速等关键声学特征。
Explaining Speaker and Spoof Embeddings via Probing
- 用简单分类器探测伪造嵌入中的说话人信息
- 在ASVspoof 2019数据集上验证出保留性别、语速、基频等特征
- 适合关注语音安全与模型可解释性的研究人员
本研究探讨基于深度神经网络的现代语音伪造检测系统中伪造嵌入(spoof embeddings)的可解释性。借鉴说话人嵌入的研究方法,我们通过训练简单神经分类器,以说话人或伪造嵌入为输入,目标标签为说话人相关属性。这些属性分为两类:基于元数据的特征(如性别、年龄)和声学特征(如基频、语速)。在ASVspoof 2019 LA评估集上的实验表明,伪造嵌入保留了性别、语速、基频和时长等关键特征。对性别和语速的进一步分析显示,伪造检测器部分保留这些特征,可能是为了确保决策过程对它们具有鲁棒性。
原文摘要 · Abstract (English)
This study investigates the explainability of embedding representations, specifically those used in modern audio spoofing detection systems based on deep neural networks, known as spoof embeddings. Building on established work in speaker embedding explainability, we examine how well these spoof embeddings capture speaker-related information. We train simple neural classifiers using either speaker or spoof embeddings as input, with speaker-related attributes as target labels. These attributes are categorized into two groups: metadata-based traits (e.g., gender, age) and acoustic traits (e.g., fundamental frequency, speaking rate). Our experiments on the ASVspoof 2019 LA evaluation set demonstrate that spoof embeddings preserve several key traits, including gender, speaking rate, F0, and duration. Further analysis of gender and speaking rate indicates that the spoofing detector partially preserves these traits, potentially to ensure the decision process remains robust against them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。