评估17种语音嵌入模型在6个数据集上的失语症检测表现,发现跨数据集泛化能力普遍较差。
Evaluating pretrained speech embedding systems for dysarthria detection across heterogenous datasets
- 采用多数据集交叉验证与零假设对比,验证结果是否显著高于随机水平
- 跨数据集准确率明显低于同数据集,最高差达15个百分点
- 提醒临床应用中避免仅用单一数据集训练测试,需关注泛化性
本文针对失语症语音检测任务,对17种公开可用的预训练语音嵌入系统进行了全面评估。由于失语症数据集通常规模小、存在录音偏差和数据不平衡问题,研究选取了涵盖相关病症的6个不同数据集,并通过多次交叉验证估计随机基准水平。为确保结果显著,将实际得分分布与精心设计的零假设分布进行比较。实验结果显示,无论使用何种嵌入模型,同数据集上的性能差异显著,质疑现有基准数据集的代表性。跨数据集测试时,准确率普遍低于同数据集,最高下降15个百分点,凸显系统泛化能力的挑战。这些发现对基于单一数据集训练的系统临床有效性提出警示。
原文摘要 · Abstract (English)
We present a comprehensive evaluation of pretrained speech embedding systems for the detection of dysarthric speech using existing accessible data. Dysarthric speech datasets are often small and can suffer from recording biases as well as data imbalance. To address these we selected a range of datasets covering related conditions and adopt the use of several cross-validations runs to estimate the chance level. To certify that results are above chance, we compare the distribution of scores across these runs against the distribution of scores of a carefully crafted null hypothesis. In this manner, we evaluate 17 publicly available speech embedding systems across 6 different datasets, reporting the cross-validation performance on each. We also report cross-dataset results derived when training with one particular dataset and testing with another. We observed that within-dataset results vary considerably depending on the dataset, regardless of the embedding used, raising questions about which datasets should be used for benchmarking. We found that cross-dataset accuracy is, as expected, lower than within-dataset, highlighting challenges in the generalization of the systems. These findings have important implications for the clinical validity of systems trained and tested on the same dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。