arXiv:2510.22225cs.CV2025-10

用语音的时频双域特征提升抑郁诊断准确率

Audio Frequency-Time Dual Domain Evaluation on Depression Diagnosis

  • 融合语音信号的时域与频域特征,构建双域分析模型
  • 在公开数据集上实现高精度抑郁分类,性能优于传统方法
  • 适合临床辅助筛查,为心理评估提供智能化新思路

抑郁症作为一种典型的精神障碍,已成为严重影响公共健康的重要问题。然而,其预防与治疗仍面临诊断流程复杂、标准模糊、就诊率低等多重挑战,严重阻碍了及时评估与干预。为此,本研究采用语音作为生理信号,利用其时频双域多模态特性,结合深度学习模型,开发了一种智能抑郁评估与诊断算法。实验结果表明,所提方法在抑郁诊断分类任务中表现优异,为抑郁的评估、筛查与诊断提供了新的视角与技术路径。

原文摘要 · Abstract (English)

Depression, as a typical mental disorder, has become a prevalent issue significantly impacting public health. However, the prevention and treatment of depression still face multiple challenges, including complex diagnostic procedures, ambiguous criteria, and low consultation rates, which severely hinder timely assessment and intervention. To address these issues, this study adopts voice as a physiological signal and leverages its frequency-time dual domain multimodal characteristics along with deep learning models to develop an intelligent assessment and diagnostic algorithm for depression. Experimental results demonstrate that the proposed method achieves excellent performance in the classification task for depression diagnosis, offering new insights and approaches for the assessment, screening, and diagnosis of depression.

抑郁诊断语音分析深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。