arXiv:2512.24628cs.SDcs.AI2025-12

用声音特征自动分类良性嗓音疾病,提升临床筛查效率。

AI-Driven Acoustic Voice Biomarker-Based Hierarchical Classification of Benign Laryngeal Voice Disorders from Sustained Vowels

  • 分三阶段构建临床导向的分级分类框架
  • 在1.26万段语音上达到94.7%准确率
  • 融合深度谱图与可解释声学特征,适合医生辅助诊断

良性喉部嗓音障碍影响近五分之一人群,常表现为发声困难,并可作为全身生理功能异常的无创指标。本文提出一种临床启发式分层机器学习框架,基于短时持续元音发音的声学特征,对八类良性嗓音障碍及健康对照进行自动分类。实验使用来自萨尔布吕肯语音数据库的15,132段录音,涵盖/a/、/i/、/u/三个元音在中、高、低及滑动音高的条件下的发音。框架分三步运行:第一阶段通过融合卷积神经网络提取的梅尔频谱特征与21个可解释的声学生物标志物,实现病理性与非病理性声音的二分类;第二阶段利用三次支持向量机将声音分为健康、功能性或精神性、结构性或炎症性三类;第三阶段结合前两阶段的概率输出,增强对结构性与炎症性疾病的区分能力,优于功能性障碍。该系统持续优于扁平多分类器及预训练自监督模型(如META HuBERT和Google HeAR),后者的目标未针对临床持续发音优化。通过结合深度谱表示与可解释声学特征,提升了模型透明度与临床契合度。结果表明,量化嗓音生物标志物具有作为可扩展、无创工具用于早期筛查、诊断分诊和长期嗓音健康管理的巨大潜力。

原文摘要 · Abstract (English)

Benign laryngeal voice disorders affect nearly one in five individuals and often manifest as dysphonia, while also serving as non-invasive indicators of broader physiological dysfunction. We introduce a clinically inspired hierarchical machine learning framework for automated classification of eight benign voice disorders alongside healthy controls, using acoustic features extracted from short, sustained vowel phonations. Experiments utilized 15,132 recordings from 1,261 speakers in the Saarbruecken Voice Database, covering vowels /a/, /i/, and /u/ at neutral, high, low, and gliding pitches. Mirroring clinical triage workflows, the framework operates in three sequential stages: Stage 1 performs binary screening of pathological versus non-pathological voices by integrating convolutional neural network-derived mel-spectrogram features with 21 interpretable acoustic biomarkers; Stage 2 stratifies voices into Healthy, Functional or Psychogenic, and Structural or Inflammatory groups using a cubic support vector machine; Stage 3 achieves fine-grained classification by incorporating probabilistic outputs from prior stages, improving discrimination of structural and inflammatory disorders relative to functional conditions. The proposed system consistently outperformed flat multi-class classifiers and pre-trained self-supervised models, including META HuBERT and Google HeAR, whose generic objectives are not optimized for sustained clinical phonation. By combining deep spectral representations with interpretable acoustic features, the framework enhances transparency and clinical alignment. These results highlight the potential of quantitative voice biomarkers as scalable, non-invasive tools for early screening, diagnostic triage, and longitudinal monitoring of vocal health.

语音分析医学诊断机器学习生物标志物

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。