用声音数据预测抑郁焦虑症状,模型可解释且公平可靠。
A multimodal Bayesian Network for symptom-level depression and anxiety prediction from voice and speech data
- 构建多模态贝叶斯网络,融合语音特征预测症状
- 在3万多名患者数据上实现症状预测AUC超0.74,误差极低
- 结果透明可解释,适合临床辅助诊断场景
精神科评估中,医生不仅关注患者自述,还重视语调、语速、流畅性等非语言信号。整合这些多源信息极具挑战,是智能工具助力的理想场景,但尚未在临床落地。本文提出基于贝叶斯网络的建模方法,利用大规模数据(30,135名独特参与者)从语音特征中预测抑郁与焦虑症状。模型在疾病层面表现优异:抑郁与焦虑的ROC-AUC分别为0.842和0.831,期望校准误差(ECE)为0.018和0.015;个体核心症状预测的ROC-AUC均超过0.74。同时评估了人口统计学公平性,并分析不同模态间的整合与冗余。探索了临床可用性及心理健康服务使用者的接受度。当输入丰富且规模足够时,该模型以症状而非疾病类别为单位进行建模,是一种原则性强、可解释、可监督的评估支持工具。
原文摘要 · Abstract (English)
During psychiatric assessment, clinicians observe not only what patients report, but important nonverbal signs such as tone, speech rate, fluency, responsiveness, and body language. Weighing and integrating these different information sources is a challenging task and a good candidate for support by intelligence-driven tools - however this is yet to be realized in the clinic. Here, we argue that several important barriers to adoption can be addressed using Bayesian network modelling. To demonstrate this, we evaluate a model for depression and anxiety symptom prediction from voice and speech features in large-scale datasets (30,135 unique speakers). Alongside performance for conditions and symptoms (for depression, anxiety ROC-AUC=0.842,0.831 ECE=0.018,0.015; core individual symptom ROC-AUC>0.74), we assess demographic fairness and investigate integration across and redundancy between different input modality types. Clinical usefulness metrics and acceptability to mental health service users are explored. When provided with sufficiently rich and large-scale multimodal data streams and specified to represent common mental conditions at the symptom rather than disorder level, such models are a principled approach for building robust assessment support tools: providing clinically-relevant outputs in a transparent and explainable format that is directly amenable to expert clinical supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。