arXiv:2502.18285cs.CL2025-02被引 8

用不确定性建模提升精神分裂谱系语音分析的诊断精度。

Uncertainty Modeling in Multimodal Speech Analysis Across the Psychosis Spectrum

  • 融合声学与语言特征,动态调整权重应对不同任务场景。
  • 降低均方根误差,F1达83%,校准误差仅4.5e-2。
  • 可识别语音标志可靠性,助力早期筛查与个性化评估。

捕捉精神分裂谱系中微妙的语音异常极具挑战,因语音模式存在固有变异,反映个体差异及临床与非临床群体症状波动性。在语音数据中建模不确定性对预测症状严重程度和提升诊断精度至关重要。精神分裂相关语音特征贯穿整个谱系,包括非临床人群。本文提出一种不确定性感知模型,整合声学与语言特征以预测症状严重度与精神病理特质。通过量化特定模态的不确定性,模型有效应对语音变异,提升预测准确性。分析了114名参与者(32例早期精神病,82例低或高偏执型特质)的语音数据,涵盖德语结构化访谈、半结构化自传任务及叙事驱动互动。模型显著降低均方根误差,实现83% F1分数,且校准误差为4.5e-2,跨任务情境表现稳健。不确定性估计增强了模型可解释性,能区分如音高变异性、流畅性中断、频谱不稳定性等语音标志的可靠性。模型根据任务结构动态调整:结构化情境侧重声学特征,非结构化情境侧重语言特征。该方法强化了精神分裂谱系研究中的早期检测、个性化评估与临床决策。

原文摘要 · Abstract (English)

Capturing subtle speech disruptions across the psychosis spectrum is challenging because of the inherent variability in speech patterns. This variability reflects individual differences and the fluctuating nature of symptoms in both clinical and non-clinical populations. Accounting for uncertainty in speech data is essential for predicting symptom severity and improving diagnostic precision. Speech disruptions characteristic of psychosis appear across the spectrum, including in non-clinical individuals. We develop an uncertainty-aware model integrating acoustic and linguistic features to predict symptom severity and psychosis-related traits. Quantifying uncertainty in specific modalities allows the model to address speech variability, improving prediction accuracy. We analyzed speech data from 114 participants, including 32 individuals with early psychosis and 82 with low or high schizotypy, collected through structured interviews, semi-structured autobiographical tasks, and narrative-driven interactions in German. The model improved prediction accuracy, reducing RMSE and achieving an F1-score of 83% with ECE = 4.5e-2, showing robust performance across different interaction contexts. Uncertainty estimation improved model interpretability by identifying reliability differences in speech markers such as pitch variability, fluency disruptions, and spectral instability. The model dynamically adjusted to task structures, weighting acoustic features more in structured settings and linguistic features in unstructured contexts. This approach strengthens early detection, personalized assessment, and clinical decision-making in psychosis-spectrum research.

语音分析精神分裂不确定性建模多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。