arXiv:2511.14801cs.SD2025-11

通过语音特征自动推断抑郁症诊断指标,实现本地化可解释分析。

IHearYou: Linking Acoustic Features to DSM-5 Depressive Behavior Indicators

  • 基于语音声学特征与DSM-5指标建立结构化关联框架。
  • 在DAIC-WOZ数据集上验证了特征与抑郁指标方向一致的关联性。
  • 系统可在普通笔记本上实时运行,适合临床可解释性研究。

抑郁症影响全球数百万人,但诊断仍依赖主观自述和访谈,难以捕捉真实行为。本文提出IHearYou,一种聚焦语音声学特征的自动化抑郁检测方法。利用家庭环境中的被动传感,该系统提取语音特征,并通过结构化关联框架将其与DSM-5(精神疾病诊断与统计手册)指标关联,特别针对重度抑郁症。系统在本地运行以保障隐私,配备持久化存储与可视化仪表盘,可在普通笔记本上实现实时处理。为确保可复现性,采用配置驱动协议,结合错误发现率(FDR)校正与按性别分层测试。应用于DAIC-WOZ数据集,结果显示特征与抑郁指标间存在方向一致的关联;基于TESS的音频流实验验证了端到端可行性。结果表明,被动语音感知可转化为可解释的DSM-5指标评分,弥合黑箱检测与临床可解释、设备端分析之间的鸿沟。

原文摘要 · Abstract (English)

Depression affects over millions people worldwide, yet diagnosis still relies on subjective self-reports and interviews that may not capture authentic behavior. We present IHearYou, an approach to automated depression detection focused on speech acoustics. Using passive sensing in household environments, IHearYou extracts voice features and links them to DSM-5 (Diagnostic and Statistical Manual of Mental Disorders) indicators through a structured Linkage Framework instantiated for Major Depressive Disorder. The system runs locally to preserve privacy and includes a persistence schema and dashboard, presenting real-time throughput on a commodity laptop. To ensure reproducibility, we define a configuration-driven protocol with False Discovery Rate (FDR) correction and gender-stratified testing. Applied to the DAIC-WOZ dataset, this protocol reveals directionally consistent feature-indicator associations, while a TESS-based audio streaming experiment validates end-to-end feasibility. Our results show how passive voice sensing can be turned into explainable DSM-5 indicator scores, bridging the gap between black-box detection and clinically interpretable, on-device analysis.

抑郁症检测语音分析可解释性本地计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。