arXiv:2506.01129cs.SDeess.AS2025-06中稿 · Interspeech 2025被引 3

比较三款语音分析工具在精神分裂症研究中的表现差异。

Comparative Evaluation of Acoustic Feature Extraction Tools for Clinical Speech Analysis

  • 统一参数后对比OpenSMILE、Praat、Librosa三工具性能
  • 基频标准差和共振峰值跨工具一致性差,甚至负相关
  • 建议临床语音研究采用多工具验证与透明报告

本研究对比了三种声学特征提取工具(OpenSMILE、Praat、Librosa)在精神分裂症谱系障碍(SSD)患者与健康对照(HC)群体的语音数据上的表现。通过标准化各工具的提取参数,分析了77名SSD患者与87名HC参与者的语音样本,发现工具间存在显著差异。尽管基频百分位数表现出高相关性(r=0.962至0.999),但基频标准差与共振峰值等指标的跨工具一致性较差,甚至出现负相关。此外,不同组别间的相关模式也存在差异。分类分析显示,基频均值、谐噪比(HNR)和MFCC1具有较高判别能力(AUC > 0.70)。研究强调了可重复性的挑战,主张建立标准化协议,开展多工具交叉验证,并加强结果透明度报告。

原文摘要 · Abstract (English)

This study compares three acoustic feature extraction toolkits (OpenSMILE, Praat, and Librosa) applied to clinical speech data from individuals with schizophrenia spectrum disorders (SSD) and healthy controls (HC). By standardizing extraction parameters across the toolkits, we analyzed speech samples from 77 SSD and 87 HC participants and found significant toolkit-dependent variations. While F0 percentiles showed high cross-toolkit correlation (r=0.962 to 0.999), measures like F0 standard deviation and formant values often had poor, even negative, agreement. Additionally, correlation patterns differed between SSD and HC groups. Classification analysis identified F0 mean, HNR, and MFCC1 (AUC greater than 0.70) as promising discriminators. These findings underscore reproducibility concerns and advocate for standardized protocols, multi-toolkit cross-validation, and transparent reporting.

语音分析精神疾病特征提取工具对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。