通过分性别层级建模,提升语音病理诊断准确率
GeHirNet: A Gender-Aware Hierarchical Model for Voice Pathology Classification
- 先识别性别特异性病征,再分性别分类疾病
- 在多源数据集上达97.63%准确率,MCC提升5%
- 适合关注语音诊断公平性与罕见病识别的研究者
基于人工智能的语音分析在疾病诊断中展现潜力,但现有分类器因性别相关声学差异及罕见疾病数据稀缺,难以精准识别特定病理。本文提出一种两阶段框架:首先利用ResNet-50对梅尔频谱图进行性别特异性病征识别,随后实施性别条件下的疾病分类。通过多尺度重采样和时间扭曲增强缓解类别不平衡问题。在整合四个公开数据源的合并数据集上,该两阶段架构结合时间扭曲增强,达到97.63%准确率和95.25%马修斯相关系数(MCC),较单阶段基线提升5% MCC。本工作推进了语音病理分类技术,通过分层建模降低性别偏差。
原文摘要 · Abstract (English)
AI-based voice analysis shows promise for disease diagnostics, but existing classifiers often fail to accurately identify specific pathologies because of gender-related acoustic variations and the scarcity of data for rare diseases. We propose a novel two-stage framework that first identifies gender-specific pathological patterns using ResNet-50 on Mel spectrograms, then performs gender-conditioned disease classification. We address class imbalance through multi-scale resampling and time warping augmentation. Evaluated on a merged dataset from four public repositories, our two-stage architecture with time warping achieves state-of-the-art performance (97.63\% accuracy, 95.25\% MCC), with a 5\% MCC improvement over single-stage baseline. This work advances voice pathology classification while reducing gender bias through hierarchical modeling of vocal characteristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。