用语音大模型统一多种测试范式,提升青少年自杀风险识别通用性。
Towards Paradigm-General Suicide Risk Detection via Speech LLM
- 构建混合专家架构,动态融合多类语音任务特征。
- 在1223人、10种范式上实现更优性能与泛化能力。
- 适合需跨场景应用的自杀风险筛查系统研发者。
青少年自杀风险仍是重大公共健康问题,语音提供了一种无创且可扩展的检测途径。现有语音风险评估通常针对单一测试范式(如言语流畅性、朗读或问答)设计,难以跨范式迁移。本文首次探索跨范式方法,将多种语音采集范式统一建模。我们以语音大模型为骨干,引入混合DoRA专家(MoDE)结构,动态捕捉不同任务间的互补线索,在涵盖1,223名参与者、10种不同语音范式的数据集上进行验证。结果表明,MoDE优于单范式专用模型和传统联合学习模型,且能有效泛化至未见范式,具备更优的置信度校准能力。
原文摘要 · Abstract (English)
Suicide risk among adolescents remains a critical public health concern, and speech provides a non-invasive and scalable approach for its detection. Speech-based suicide risk assessment commonly relies on carefully designed speech elicitation paradigms (\textit{e.g.,} verbal fluency, reading, or question answering) to probe cognitive and affective states. Existing approaches, however, typically focus on one single paradigm at a time. This paper, for the first time, investigates cross-paradigm approaches that unify diverse speech elicitation paradigms within a single model. Specifically, we use a speech LLM as backbone with a mixture of DoRA experts (MoDE) to capture complementary cues across assessments dynamically, tested on 1,223 participants across ten speech elicitation paradigms. Results show that MoDE outperforms both paradigm-specific and conventional joint-learning models. Moreover, it can generalise to unseen paradigms and provide better confidence calibration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。