arXiv:2505.13069cs.CLcs.LG2025-05被引 1

用多模态语音特征评估青少年自杀风险,融合文本与音频信息提升识别准确率。

Suicide Risk Assessment Using Multimodal Speech Features: A Study on the SW1 Challenge Dataset

  • 融合语音转录、语言与音频嵌入,结合手工提取的声学特征
  • 加权注意力融合策略在开发集上达到69%准确率,但测试集表现有差距
  • 适合关注心理健康监测与多模态机器学习的研究者

首届SpeechWellness挑战赛凸显了针对青少年进行基于语音的自杀风险评估的重要性。本研究探索了一种多模态方法,整合WhisperX自动转录、中文RoBERTa语言嵌入和WavLM音频嵌入,并引入了MFCC、谱对比度及与音高相关的统计特征等手工声学特征。比较了三种融合策略:早期拼接、模态特异性处理以及带mixup正则化的加权注意力机制。结果显示,加权注意力策略泛化能力最佳,在开发集上达到69%准确率,但开发集与测试集间存在性能差距,反映出泛化挑战。研究严格基于MINI-KID框架,强调需优化嵌入表示与融合机制以提高分类可靠性。

原文摘要 · Abstract (English)

The 1st SpeechWellness Challenge conveys the need for speech-based suicide risk assessment in adolescents. This study investigates a multimodal approach for this challenge, integrating automatic transcription with WhisperX, linguistic embeddings from Chinese RoBERTa, and audio embeddings from WavLM. Additionally, handcrafted acoustic features -- including MFCCs, spectral contrast, and pitch-related statistics -- were incorporated. We explored three fusion strategies: early concatenation, modality-specific processing, and weighted attention with mixup regularization. Results show that weighted attention provided the best generalization, achieving 69% accuracy on the development set, though a performance gap between development and test sets highlights generalization challenges. Our findings, strictly tied to the MINI-KID framework, emphasize the importance of refining embedding representations and fusion mechanisms to enhance classification reliability.

自杀风险评估多模态语音分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。