用量子视觉理论提升语音伪造检测准确率
Quantum Vision Theory Applied to Audio Classification for Deepfake Speech Detection
- 将语音谱图转为信息波再输入模型,突破传统表示方式
- 在ASVspoof数据集上最高达94.57%准确率,误报率仅9.04%
- 适合做语音安全与深度伪造检测的研究者参考
我们提出量子视觉(QV)理论作为基于深度学习的音频分类新视角,应用于深度伪造语音检测。受量子物理中波粒二象性启发,QV理论认为数据不仅可表现为可观测的“坍缩”形式,还可作为信息波存在。传统深度学习直接使用图像等坍缩表示进行训练,而QV理论先通过QV模块将输入转化为信息波,再送入深度模型。在图像分类任务中,基于QV的模型表现优于传统方法。本研究将该理论扩展至语音领域:利用短时傅里叶变换(STFT)、梅尔谱图和梅尔频率倒谱系数(MFCC),通过所提出的QV模块生成信息波,用于训练基于QV的卷积神经网络(QV-CNN)与视觉变换器(QV-ViT)。在ASVSpoof数据集上进行了大量实验,结果表明,QV-CNN与QV-ViT持续优于标准CNN与ViT模型,在区分真实与伪造语音方面表现出更高准确率与鲁棒性。其中,基于MFCC特征的QV-CNN达到94.20%准确率与9.04%等错误率(EER),而基于梅尔谱图的QV-CNN取得最高94.57%准确率。这些发现验证了QV理论在语音深度伪造检测中的有效性与前景,为音频感知任务中量子启发学习开辟了新方向。
原文摘要 · Abstract (English)
We propose Quantum Vision (QV) theory as a new perspective for deep learning-based audio classification, applied to deepfake speech detection. Inspired by particle-wave duality in quantum physics, QV theory is based on the idea that data can be represented not only in its observable, collapsed form, but also as information waves. In conventional deep learning, models are trained directly on these collapsed representations, such as images. In QV theory, inputs are first transformed into information waves using a QV block, and then fed into deep learning models for classification. QV-based models improve performance in image classification compared to their non-QV counterparts. What if QV theory is applied speech spectrograms for audio classification tasks? This is the motivation and novelty of the proposed approach. In this work, Short-Time Fourier Transform (STFT), Mel-spectrograms, and Mel-Frequency Cepstral Coefficients (MFCC) of speech signals are converted into information waves using the proposed QV block and used to train QV-based Convolutional Neural Networks (QV-CNN) and QV-based Vision Transformers (QV-ViT). Extensive experiments are conducted on the ASVSpoof dataset for deepfake speech classification. The results show that QV-CNN and QV-ViT consistently outperform standard CNN and ViT models, achieving higher classification accuracy and improved robustness in distinguishing genuine and spoofed speech. Moreover, the QV-CNN model using MFCC features achieves the best overall performance on the ASVspoof dataset, with an accuracy of 94.20% and an EER of 9.04%, while the QV-CNN with Mel-spectrograms attains the highest accuracy of 94.57%. These findings demonstrate that QV theory is an effective and promising approach for audio deepfake detection and opens new directions for quantum-inspired learning in audio perception tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。