通过向量对齐实现跨语言发音障碍检测,提升帕金森病语音识别准确率。
Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
- 用健康人群语音估计中心点,对齐源语言与目标语言的语音表征。
- 在捷克、德语、西班牙语数据上,跨语言检测敏感性与F1显著提升。
- 适用于缺乏标注数据的跨语言医疗语音分析,尤其适合多语言研究者。
发音障碍语音数据稀缺,使得跨语言检测成为重要但具挑战性的问题。一个关键难点是语音表征常包含语言特异性结构,干扰发音障碍识别。本文提出一种表征级语言迁移(LS)方法,利用健康对照组语音估算的中心点向量,将源语言自监督语音表征对齐至目标语言分布。我们在帕金森病语音数据集中的捷克语、德语和西班牙语口述DDK录音上,评估了该方法在跨语言与多语言设置下的表现。结果表明,LS在跨语言设置下显著提升灵敏度与F1值,多语言设置中也获得稳定小幅增益。表征分析进一步显示,LS降低了嵌入空间中的语言标识性,支持其去除语言依赖结构的解释。
原文摘要 · Abstract (English)
The limited availability of dysarthric speech data makes cross-lingual detection an important but challenging problem. A key difficulty is that speech representations often encode language-dependent structure that can confound dysarthria detection. We propose a representation-level language shift (LS) that aligns source-language self-supervised speech representations with the target-language distribution using centroid-based vector adaptation estimated from healthy-control speech. We evaluate the approach on oral DDK recordings from Parkinson's disease speech datasets in Czech, German, and Spanish under both cross-lingual and multilingual settings. LS substantially improves sensitivity and F1 in cross-lingual settings, while yielding smaller but consistent gains in multilingual settings. Representation analysis further shows that LS reduces language identity in the embedding space, supporting the interpretation that LS removes language-dependent structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。