用频谱相关性捕捉语音周期性特征,提升深度伪造语音检测效果
Cyclostationarity Analysis as a Complement to Self-Supervised Representations for Speech Deepfake Detection
- 基于频谱相关密度构建周期性统计特征,挖掘语音高频结构
- 在ASVspoof 2019 LA上将误检率从8.28%降至0.98%
- 适合关注信号级特征与自监督模型融合的语音安全研究者
语音深度伪造检测(SDD)对维护语音驱动技术的信任至关重要。尽管当前SDD系统越来越多依赖能捕捉丰富上下文信息的自监督学习(SSL)表示,但信号驱动的声学特征仍对建模语音细粒度结构特性具有重要意义。现有声学前端多基于时频表示,未能充分挖掘语音信号中的高阶谱依赖关系。本文提出一种基于谱相关密度(SCD)的周期平稳性启发式特征提取框架,通过捕捉频率分量间的谱相关性,建模语音的周期性统计结构。特别地,研究了周期平稳特征是否能为传统时频特征和现代SSL嵌入提供互补信息。为此,引入时序结构化的SCD表示,保留谱相关性和循环频率相关性随时间的变化。在ASVspoof 2019 LA、2021 DF和ASVspoof 5数据集上评估多种SDD架构,结果表明高阶谱相关性在多个挑战场景下提供了有效互补信息。尤其在ASVspoof 2019 LA上,融合SSL与SCD嵌入使等错误率从8.28%降至0.98%;在更具挑战性的ASVspoof 5数据集上,从15.81%降至14.80%。这些发现确立了高阶周期平稳谱相关性作为现代SDD的互补信息源,同时为自然与合成语音之间的统计差异提供了新的信号处理洞察。
原文摘要 · Abstract (English)
Speech deepfake detection (SDD) is essential for maintaining trust in voice-driven technologies and digital media. Although recent SDD systems increasingly rely on SSL representations that capture rich contextual information, complementary signal-driven acoustic features remain important for modeling fine-grained structural properties of speech. Most existing acoustic front ends are based on time-frequency representations, which do not fully exploit higher-order spectral dependencies inherent in speech signals. We introduce a cyclostationarity-inspired acoustic feature extraction framework for SDD based on spectral correlation density (SCD). The proposed features model periodic statistical structures in speech by capturing spectral correlations between frequency components. In particular, we investigate whether cyclostationary representations provide complementary information beyond conventional time-frequency features and modern SSL embeddings. To this end, we introduce temporally structured SCD representations that preserve the evolution of spectral and cyclic-frequency correlations over time. Their effectiveness is evaluated using multiple SDD architectures, including conventional neural networks, SSL-based embedding systems, and hybrid fusion models. Experiments on ASVspoof 2019 LA, 2021 DF, and ASVspoof 5 demonstrate that higher-order spectral correlations provide complementary information to SSL embeddings in several challenging conditions. In particular, fusion of SSL and SCD embeddings reduces the equal error rate on ASVspoof 2019 LA from 8.28% to 0.98%, and from 15.81% to 14.80% on the challenging ASVspoof 5 dataset. These findings establish higher-order cyclostationary spectral correlations as a complementary source of information for modern SDD, while providing new signal-processing insight into the statistical differences between natural and synthetic speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。