arXiv:2503.10301eess.AScs.AI2025-03中稿 · ICASSP 2025 - Pers…被引 17

用双头模型同时分析发音和口语,提升跨语言帕金森病语音检测效果

Bilingual Dual-Head Deep Model for Parkinson's Disease Detection from Speech

  • 双头结构分别处理节奏性发音和自然口语,按输入类型自动切换
  • 在斯洛伐克语和西班牙语数据集上,跨语言检测准确率显著提升
  • 适合需要多语言医疗语音分析的研究者与临床应用开发者

本文针对多语言环境下帕金森病语音检测问题,提出一种专用的双头深度神经网络架构,用于基于类型的二分类任务。一个分支专用于分析节律性发音(diadochokinetic)模式,另一个分支则捕捉连续话语中的自然语音特征。根据输入内容自动启用对应分支。语音表示通过自监督学习(SSL)模型与小波变换提取,结合自适应层、卷积瓶颈和对比学习以减少语言间差异。在覆盖斯洛伐克语的EWA-DB和西班牙语的PC-GITA两个数据集上进行评估。结果表明,单一语言训练的模型难以实现跨语言泛化,简单拼接数据集效果也不佳。而本模型在两种语言上均实现良好泛化性能。

原文摘要 · Abstract (English)

This work aims to tackle the Parkinson's disease (PD) detection problem from the speech signal in a bilingual setting by proposing an ad-hoc dual-head deep neural architecture for type-based binary classification. One head is specialized for diadochokinetic patterns. The other head looks for natural speech patterns present in continuous spoken utterances. Only one of the two heads is operative accordingly to the nature of the input. Speech representations are extracted from self-supervised learning (SSL) models and wavelet transforms. Adaptive layers, convolutional bottlenecks, and contrastive learning are exploited to reduce variations across languages. Our solution is assessed against two distinct datasets, EWA-DB, and PC-GITA, which cover Slovak and Spanish languages, respectively. Results indicate that conventional models trained on a single language dataset struggle with cross-linguistic generalization, and naive combinations of datasets are suboptimal. In contrast, our model improves generalization on both languages, simultaneously.

帕金森病语音分析双头模型多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。