用30秒数数声纹预测多种健康状况,效果优于传统方法
HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification
- 用30秒数数语音数据,测试14种声纹模型的健康预测能力
- 男性的声纹模型对中重度睡眠呼吸暂停预测AUC达0.64,优于MFCC和人口统计
- 不同性别对不同模型效果差异明显,指导未来临床应用选择
人类语音包含反映生理与神经状态的副语言线索,可能实现多种医学表型的无创检测。我们提出人类表型项目语音数据集(HPP-Voice):包含7,188段希伯来语成年者30秒数数录音,每名参与者关联最多15种与声音相关的表型,涵盖呼吸、睡眠、心理健康、代谢、免疫及神经系统疾病。我们系统比较了14种现代语音嵌入模型,发现这些30秒数数任务生成的嵌入在下游健康分类中表现优于MFCC和人口统计特征。基于说话人识别模型学习的嵌入,在男性中可预测客观测量的中重度睡眠呼吸暂停,AUC为0.64 ± 0.03;而MFCC和人口统计特征分别仅为0.56 ± 0.02和0.57 ± 0.02。此外,结果揭示不同医疗领域存在性别特异性模型效能差异:男性中,说话人识别与语音分段模型在呼吸类(如哮喘:0.61 ± 0.03 vs. 0.56 ± 0.02)和睡眠相关疾病(如失眠:0.65 ± 0.04 vs. 0.59 ± 0.05)上优于语音基础模型;女性中,语音分段模型在吸烟状态分类中最佳(0.61 ± 0.02 vs. 0.55 ± 0.02),而希伯来语专用模型在焦虑分类中略优(0.59 ± 0.02 vs. 0.58 ± 0.02)。研究证明简单数数任务可支持大规模多表型语音筛查,并指明各类嵌入模型在特定条件下的泛化优势,为未来声学生物标志物研究与临床部署提供依据。
原文摘要 · Abstract (English)
Human speech contains paralinguistic cues that reflect a speaker's physiological and neurological state, potentially enabling non-invasive detection of various medical phenotypes. We introduce the Human Phenotype Project Voice corpus (HPP-Voice): a dataset of 7,188 recordings in which Hebrew-speaking adults count for 30 seconds, with each speaker linked to up to 15 potentially voice-related phenotypes spanning respiratory, sleep, mental health, metabolic, immune, and neurological conditions. We present a systematic comparison of 14 modern speech embedding models, where modern speech embeddings from these 30-second counting tasks outperform MFCCs and demographics for downstream health condition classifications. We found that embedding learned from a speaker identification model can predict objectively measured moderate to severe sleep apnea in males with an AUC of 0.64 $\pm$ 0.03, while MFCC and demographic features led to AUCs of 0.56 $\pm$ 0.02 and 0.57 $\pm$ 0.02, respectively. Additionally, our results reveal gender-specific patterns in model effectiveness across different medical domains. For males, speaker identification and diarization models consistently outperformed speech foundation models for respiratory conditions (e.g., asthma: 0.61 $\pm$ 0.03 vs. 0.56 $\pm$ 0.02) and sleep-related conditions (insomnia: 0.65 $\pm$ 0.04 vs. 0.59 $\pm$ 0.05). For females, speaker diarization models performed best for smoking status (0.61 $\pm$ 0.02 vs 0.55 $\pm$ 0.02), while Hebrew-specific models performed best (0.59 $\pm$ 0.02 vs. 0.58 $\pm$ 0.02) in classifying anxiety compared to speech foundation models. Our findings provide evidence that a simple counting task can support large-scale, multi-phenotypic voice screening and highlight which embedding families generalize best to specific conditions, insights that can guide future vocal biomarker research and clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。