用多任务变压器检测语音伪造,还能解释判断依据。
Multi-Task Transformer for Explainable Speech Deepfake Detection via Formant Modeling
- 同时预测共振峰轨迹和声带振动模式
- 比基线模型参数更少、训练更快、准确率不降
- 能定位判断依赖的有声/无声区域,适合安全审核
本文提出一种用于语音深度伪造检测的多任务变换器,可随时间预测共振峰轨迹和声带振动模式,最终分类语音为真实或伪造,并指出其决策更依赖于有声或无声区域。在先前的说话人-共振峰变换器架构基础上,我们通过优化输入分段策略、重构解码过程并集成内置可解释性,使模型参数更少、训练速度更快,且保持预测性能,同时提升可解释性。
原文摘要 · Abstract (English)
In this work, we introduce a multi-task transformer for speech deepfake detection, capable of predicting formant trajectories and voicing patterns over time, ultimately classifying speech as real or fake, and highlighting whether its decisions rely more on voiced or unvoiced regions. Building on a prior speaker-formant transformer architecture, we streamline the model with an improved input segmentation strategy, redesign the decoding process, and integrate built-in explainability. Compared to the baseline, our model requires fewer parameters, trains faster, and provides better interpretability, without sacrificing prediction performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。