arXiv:2412.16900eess.AScs.CL2024-12被引 26

仅迁移编码器权重,用海量语音数据提升抑郁症预测准确率

Speech-Based Depression Prediction Using Encoder-Weight-Only Transfer Learning and a Large Corpus

  • 只迁移编码器权重,降低模型运行开销
  • 相比基准提升27%准确率,统计显著(p值接近0)
  • 适合资源有限但需高效部署的临床辅助系统

基于语音的算法在管理抑郁症等行为健康问题中受到关注。本文提出一种轻量级编码器的语音迁移学习方法,仅迁移编码器权重,实现简化推理模型。研究使用包含约两数量级更多说话人和会话的大规模数据集,显著优于以往工作。实验表明,预测PHQ-8评分时二分类任务性能最高提升27%,且具有极低的p值(接近零),回归任务也获得提升。值得注意的是,迁移学习收益不依赖源任务的强性能表现。结果表明该方法灵活高效,具备实际部署潜力。

原文摘要 · Abstract (English)

Speech-based algorithms have gained interest for the management of behavioral health conditions such as depression. We explore a speech-based transfer learning approach that uses a lightweight encoder and that transfers only the encoder weights, enabling a simplified run-time model. Our study uses a large data set containing roughly two orders of magnitude more speakers and sessions than used in prior work. The large data set enables reliable estimation of improvement from transfer learning. Results for the prediction of PHQ-8 labels show up to 27% relative performance gains for binary classification; these gains are statistically significant with a p-value close to zero. Improvements were also found for regression. Additionally, the gain from transfer learning does not appear to require strong source task performance. Results suggest that this approach is flexible and offers promise for efficient implementation.

抑郁症预测迁移学习语音分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。