通过分层推理提升语音伪造检测效率,参数减少10.8%仍保持高性能。
REIMU: Efficient Heterogeneous Hierarchical Reasoning for SSL-Based Speech Deepfake Detection

- 采用异构分层推理架构,结合自注意力与线性注意力模块。
- 在ASVspoof 2019/2021数据集上实现高精度检测,参数量减少10.8%。
- 适合追求轻量化、高效能语音伪造检测的工业部署场景。
文本到语音和语音转换系统生成语音的真实度不断提升,对媒体完整性和语音认证构成严峻挑战。自监督学习(SSL)显著推动了语音伪造检测的发展,传统下游模型通常对SSL表征进行单次前向传播处理。本文系统研究了递归分层推理在该任务中的实际效果,提出名为REIMU的可控实验框架,对比了单次前向、权重共享递归、同构分层推理模块(HRM)与异构分层推理模块(HRM),涵盖四种基础规模的SSL前端。进一步考察了融合自注意力与线性注意力的高低层异构模块。在ASVspoof 2019和2021评测集上的实验表明,递归与分层分解本身并不必然提升性能,而异构模块设计更具竞争力。值得注意的是,该设计在保持性能的同时,比基准模型减少10.8%的下游参数量,展现出参数高效的潜力。
原文摘要 · Abstract (English)
The increasing realism of speech generated by text-to-speech and voice conversion systems poses growing challenges to media integrity and voice authentication. Self-supervised learning (SSL) has substantially advanced speech deepfake detection, where downstream backbones conventionally process SSL representations through a single forward pass. This work investigates the practical effectiveness of recurrent hierarchical reasoning for this task. We term this controlled study REIMU and systematically compare conventional single-pass backbones, weight-shared recurrence, homogeneous HRM, and heterogeneous HRM across four Base-scale SSL frontends. We further examine heterogeneous high- and low-level modules that combine self-attention with linear attention. Experiments on the ASVspoof 2019 and 2021 evaluation sets show that recurrence and hierarchical decomposition do not inherently improve detection, whereas heterogeneous operator assignment provides a more competitive configuration. Notably, the heterogeneous design remains competitive while using 10.8\% fewer downstream parameters than the matched baseline, demonstrating its potential for parameter-efficient speech deepfake detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。