提出四指标框架,检测语音中病理、情绪和语言特征的几何分离性。
Geometric Analysis of Speech Representation Spaces: Topological Disentanglement and Confound Detection
- 用四个聚类指标分析六语料八数据集的语音特征分离度。
- 情绪特征最紧凑(轮廓系数0.250),病理次之(0.141),语言最松散(0.077)。
- 发现病理与语言混杂低于0.21,适合临床部署,指导多语言语音健康系统设计。
基于语音的临床工具越来越多地应用于多语言环境,但病理语音特征是否仍能与口音差异在几何上保持分离尚不明确。系统可能误判非母语健康者或漏诊多语言患者。我们提出一个四指标聚类框架,评估六语料八数据集组合中情感、语言和病理语音特征的几何解耦程度。结果显示稳定层级:情感特征形成最紧密聚类(轮廓系数0.250),其次为病理特征(0.141),语言特征最松散(0.077)。共混分析表明,病理与语言特征重叠低于0.21,高于置换零模型但有界,适用于临床部署。可信度分析确认嵌入表示的保真性与几何结论的鲁棒性。该框架为跨多元人群的公平可靠语音健康系统提供可操作指南。
原文摘要 · Abstract (English)
Speech-based clinical tools are increasingly deployed in multilingual settings, yet whether pathological speech markers remain geometrically separable from accent variation remains unclear. Systems may misclassify healthy non-native speakers or miss pathology in multilingual patients. We propose a four-metric clustering framework to evaluate geometric disentanglement of emotional, linguistic, and pathological speech features across six corpora and eight dataset combinations. A consistent hierarchy emerges: emotional features form the tightest clusters (Silhouette 0.250), followed by pathological (0.141) and linguistic (0.077). Confound analysis shows pathological-linguistic overlap remains below 0.21, which is above the permutation null but bounded for clinical deployment. Trustworthiness analysis confirms embedding fidelity and robustness of the geometric conclusions. Our framework provides actionable guidelines for equitable and reliable speech health systems across diverse populations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。