arXiv:2604.26676cs.SDcs.AI2026-04

发现语音数据中的虚假相关性,避免模型误判。

A Toolkit for Detecting Spurious Correlations in Speech Datasets

论文配图:A Toolkit for Detecting Spurious Correlations in Speech Datasets
图 1 · 摘自论文原文
  • 通过非语音段检测标签,识别数据中的虚假关联
  • 非语音段分类性能优于随机,说明存在虚假相关
  • 适合医疗语音等高风险场景的模型验证

我们提出一个工具包,用于揭示语音数据中记录特征与目标类别之间的虚假相关性。这类相关性常出现在健康相关数据集中,因录音条件差异而产生。当训练和测试数据均存在此类相关性时,会显著高估系统性能,尤其在对性能有最低要求的高风险应用中构成严重隐患。该工具包基于一种诊断方法:仅利用音频中的非语音区域进行目标类别预测。若该任务表现优于随机水平,即表明非语音部分隐含了类别信息,从而提示虚假相关性的存在。工具包已公开供研究使用。

原文摘要 · Abstract (English)

We introduce a toolkit for uncovering spurious correlations between recording characteristics and target class in speech datasets. Spurious correlations may arise due to heterogeneous recording conditions, a common scenario for health-related datasets. When present both in the training and test data, these correlations result in an overestimation of the system performance -- a dangerous situation, specially in high-stakes application where systems are required to satisfy minimum performance requirements. Our toolkit implements a diagnostic method based on the detection of the target class using only the non-speech regions in the audio. Better than chance performance at this task indicates that information about the target class can be extracted from the non-speech regions, flagging the presence of spurious correlations. The toolkit is publicly available for research use.

语音识别虚假相关数据诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。