用多视图相关分析消除语音中的无关信息,提升病理语音检测准确率。
Multiview Canonical Correlation Analysis for Automatic Pathological Speech Detection
- 对多种输入表示应用多视图典型相关分析,剔除无关噪声。
- 相比其他降维方法,病理语音识别准确率显著提升。
- 适合关注可解释性与传统模型优化的研究者。
近期提出的自动病理语音检测方法依赖于频谱图输入或wav2vec2嵌入。这些表示可能包含与病理无关的不相关信息,如语音内容变化或说话风格的时间差异,从而影响分类性能。为解决此问题,我们提出在自动病理语音检测前对这些输入表示使用多视图典型相关分析(MCCA)。结果表明,与其它降维技术不同,MCCA能有效消除输入表示中的不相关信息,显著提升病理语音检测性能。结合传统分类器使用MCCA,性能可媲美甚至优于复杂架构,同时保持表示结构并增强可解释性。
原文摘要 · Abstract (English)
Recently proposed automatic pathological speech detection approaches rely on spectrogram input representations or wav2vec2 embeddings. These representations may contain pathology irrelevant uncorrelated information, such as changing phonetic content or variations in speaking style across time, which can adversely affect classification performance. To address this issue, we propose to use Multiview Canonical Correlation Analysis (MCCA) on these input representations prior to automatic pathological speech detection. Our results demonstrate that unlike other dimensionality reduction techniques, the use of MCCA leads to a considerable improvement in pathological speech detection performance by eliminating uncorrelated information present in the input representations. Employing MCCA with traditional classifiers yields a comparable or higher performance than using sophisticated architectures, while preserving the representation structure and providing interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。