融合语音与文本,动态加权提升自杀风险检测精度
Dynamic Fusion Multimodal Network for SpeechWellness Detection
- 采用动态加权机制融合时域、时频域语音与语义特征
- 参数减少78%,准确率提升5%,超越基准模型
- 适合心理健康监测、轻量化多模态系统研究者
自杀是青少年死亡的主要原因之一。以往研究多单独分析文本或语音信息,而融合语音与文本等多模态信号能更全面评估个体心理状态。受此启发,并针对首届SpeechWellness检测挑战赛,我们提出一种基于动态融合机制的轻量级多分支多模态系统。为突破传统仅用时域波形分析语音的局限,系统同时引入时域与时频(TF)域语音特征及语义表示。通过设计可学习权重的动态融合模块,自适应调整各模态贡献。为提升效率,对原基线模型进行轻量化简化。实验表明,该系统性能优于基准模型,参数量减少78%,准确率提升5%。
原文摘要 · Abstract (English)
Suicide is one of the leading causes of death among adolescents. Previous suicide risk prediction studies have primarily focused on either textual or acoustic information in isolation, the integration of multimodal signals, such as speech and text, offers a more comprehensive understanding of an individual's mental state. Motivated by this, and in the context of the 1st SpeechWellness detection challenge, we explore a lightweight multi-branch multimodal system based on a dynamic fusion mechanism for speechwellness detection. To address the limitation of prior approaches that rely on time-domain waveforms for acoustic analysis, our system incorporates both time-domain and time-frequency (TF) domain acoustic features, as well as semantic representations. In addition, we introduce a dynamic fusion block to adaptively integrate information from different modalities. Specifically, it applies learnable weights to each modality during the fusion process, enabling the model to adjust the contribution of each modality. To enhance computational efficiency, we design a lightweight structure by simplifying the original baseline model. Experimental results demonstrate that the proposed system exhibits superior performance compared to the challenge baseline, achieving a 78% reduction in model parameters and a 5% improvement in accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。