arXiv:2601.02432cs.SDcs.LG2026-01

量子卷积网络在语音医疗中抗噪声能力优于传统模型,尤其对时移、变调等畸变更鲁棒。

Quantifying Quanvolutional Neural Networks Robustness for Speech in Healthcare Applications

  • 用量子卷积网络(QNN)与经典CNN对比,测试其在四种噪声下的表现
  • 在时移、变调等情况下QNN误差降低最高达22%,收敛速度更快6倍
  • 适合关注语音病理和情绪识别的医疗AI研究者,尤其关注抗畸变能力

基于语音的机器学习系统易受噪声干扰,影响情感识别与语音病理检测的可靠性。本文在干净训练/噪声测试条件下,评估了混合量子机器学习模型——量子卷积神经网络(QNN)相对于经典卷积神经网络(CNN)在四种声学畸变(高斯噪声、音高偏移、时间偏移、速度变化)下的鲁棒性。使用AVFAD(语音病理)和TESS(语音情感)数据集,对比三种QNN模型(Random、Basic、Strongly)与CNN-Base、ResNet-18、VGG-16的准确率及腐蚀度量(CE、mCE、RCE、RmCE),分析电路复杂度、收敛性及各情感类别鲁棒性。结果显示,QNN在音高偏移、时间偏移和速度变化下普遍优于CNN-Base(严重时间偏移时CE/RCE低至22%),而CNN-Base在高斯噪声下仍更稳健。QNN-Basic在AVFAD上整体表现最佳,QNN-Random在TESS上最优。情感层面,恐惧最鲁棒(严重畸变下准确率80-90%),中性在强高斯噪声下崩溃(准确率仅5.5%),快乐对音高、时间、速度畸变最敏感。此外,QNN收敛速度比CNN-Base快至六倍。据我们所知,这是首个系统研究QNN在常见非对抗性声学畸变下语音鲁棒性的工作,表明浅层纠缠量子前处理可提升抗噪能力,但对加性噪声的敏感性仍是挑战。

原文摘要 · Abstract (English)

Speech-based machine learning systems are sensitive to noise, complicating reliable deployment in emotion recognition and voice pathology detection. We evaluate the robustness of a hybrid quantum machine learning model, quanvolutional neural networks (QNNs) against classical convolutional neural networks (CNNs) under four acoustic corruptions (Gaussian noise, pitch shift, temporal shift, and speed variation) in a clean-train/corrupted-test regime. Using AVFAD (voice pathology) and TESS (speech emotion), we compare three QNN models (Random, Basic, Strongly) to a simple CNN baseline (CNN-Base), ResNet-18 and VGG-16 using accuracy and corruption metrics (CE, mCE, RCE, RmCE), and analyze architectural factors (circuit complexity or depth, convergence) alongside per-emotion robustness. QNNs generally outperform the CNN-Base under pitch shift, temporal shift, and speed variation (up to 22% lower CE/RCE at severe temporal shift), while the CNN-Base remains more resilient to Gaussian noise. Among quantum circuits, QNN-Basic achieves the best overall robustness on AVFAD, and QNN-Random performs strongest on TESS. Emotion-wise, fear is most robust (80-90% accuracy under severe corruptions), neutral can collapse under strong Gaussian noise (5.5% accuracy), and happy is most vulnerable to pitch, temporal, and speed distortions. QNNs also converge up to six times faster than the CNN-Base. To our knowledge, this is a systematic study of QNN robustness for speech under common non-adversarial acoustic corruptions, indicating that shallow entangling quantum front-ends can improve noise resilience while sensitivity to additive noise remains a challenge.

量子机器学习语音病理鲁棒性医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。