arXiv:2606.22910cs.SDcs.AI2026-06中稿 · Interspeech 2026

用跨语言语音增强病理语音分类,解决标注数据少的问题。

Cross-lingual Retrieval-Augmented Classification for Dysarthria Severity Assessment

论文配图:Cross-lingual Retrieval-Augmented Classification for Dysarthria Severity Assessment
图 1 · 摘自论文原文
  • 通过跨语言检索融合,利用异语言语音提升分类性能。
  • 韩语和意大利语数据集上准确率分别达87.3%和86.7%,提升超20个百分点。
  • 适合缺乏标注数据的言语障碍评估场景,尤其适用于多语言研究。

自动发音障碍严重程度评估受限于病理语音标注数据稀缺。为此,我们提出跨语言检索增强分类(CRAC),通过异语言语音构建辅助信号。首先采用监督对比学习构建聚焦严重程度的嵌入空间,再从异语言语料库中建立向量数据库。训练与推理时,分类器在对齐空间中检索前k个参考样本,并通过交叉注意力将其与输入融合。在韩语卒中后及意大利肌萎缩侧索硬化症(ALS)发音障碍数据集上,采用说话人无关三分类协议评估,CRAC分别取得87.3%和86.7%的平衡准确率,相较单语言基线分别提升8.4和20.0个百分点。

原文摘要 · Abstract (English)

Automatic dysarthria severity assessment is limited by the scarcity of labeled pathological speech data. To address this, we propose Cross-lingual Retrieval-Augmented Classification (CRAC), which leverages speech from a different language via an align-retrieve-fuse pipeline. Supervised contrastive learning first shapes a severity-focused embedding space, then a vector database is built from the opposite-language corpus. During both training and inference, the classifier retrieves top-k references from the aligned space and fuses them with the input via cross-attention. Evaluated on Korean post-stroke and Italian ALS dysarthria datasets under a speaker-independent three-class protocol, CRAC achieves balanced accuracies of 87.3% on Korean and 86.7% on Italian, improving over monolingual baselines by 8.4 and 20.0 percentage points, respectively.

语音分析跨语言医学评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。