构建3.2小时西班牙语神经疾病语音数据集,助力语音识别研究
S-DiverSe: Spanish Diverse Speech

- 收集22名患者真实场景语音,含病种与可懂度标注
- 基线识别错误率高达45.8%,凸显现有模型短板
- 文本后处理比微调更有效,适合临床实用场景
自动语音识别(ASR)在标准语音上已取得显著进展,但受神经系统疾病影响的语音仍具挑战。本文提出S-DiverSe(西班牙语多样语音)数据集,包含22名患有肌萎缩侧索硬化症、帕金森病和中风患者的3.2小时真实环境西班牙语语音,共444段人工转录音频,并附有说话人性别、疾病类型及可懂度等元数据。该数据集旨在支持神经疾病相关西班牙语语音的识别评估与模型开发。我们描述了数据集构成,分析其分布特征,并报告基线ASR结果及初步适应实验。结果显示,对于域外神经疾病西班牙语语音,启发式文本后处理比微调更具鲁棒性。这凸显了构建专用真实场景西班牙语基准的重要性。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) has advanced remarkably for standard speech, yet speech affected by neurological conditions remains a challenge. We present S-DiverSe (Spanish Diverse Speech), a corpus of 3.2 hours of in-the-wild Spanish speech from 22 speakers with amyotrophic lateral sclerosis, Parkinson's disease, and stroke. The dataset contains 444 manually transcribed audio segments with metadata on speaker sex, disease type, and intelligibility. S-DiverSe is designed to support ASR evaluation and development for neurologically affected Spanish speech. We describe the dataset, analyze its composition, and report baseline ASR results alongside initial adaptation experiments. Our findings reveal that heuristic text post-processing is more robust than fine-tuning for out-of-domain neurological Spanish speech. This underscores the need for dedicated in-the-wild Spanish benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。