用对话节奏提升双人抑郁检测,效果优于传统语音分析。
Can Conversational Temporal Dynamics Improve Depression Detection in Dyads? A Preliminary Investigation in Multi-Modality Perspectives

- 引入双人对话轮次时间作为新模态,与自监督模型融合。
- 在DAIC-WOZ数据集上达0.804和0.669的宏平均F1,优于基线。
- 模型自动忽略声学特征,凸显对话时序的解释性优势。
从临床访谈中自动识别抑郁通常聚焦于参与者的语义内容和声学特征,但医生与患者之间的互动时序仍较少被建模。本文研究对话时序动态,特别是双人轮次时间,作为主要模态与自监督编码器融合。在DAIC-WOZ数据集上,对比一个24维紧凑时序模块与冻结的WavLM-large和RoBERTa-large基线检测器。该时序模块在开发集上取得最高单模态性能。此外,凸加权后融合策略使整体性能提升至开发集0.804、测试集0.669的宏平均F1。学习到的融合权重将声学特征置为零,表明对话时序是双人抑郁筛查中轻量且可解释的补充模态。
原文摘要 · Abstract (English)
Automatic depression detection from clinical interviews typically models the semantic content and acoustic characteristics of participant speech. However, the interactional timing between the clinician and participant remains comparatively under-modeled. We investigate conversational temporal dynamics, specifically dyadic turn-pair timing, as a primary modality fused with self-supervised encoders. Evaluated on the DAIC-WOZ dataset, we compare a compact 24-dimensional timing module against frozen WavLM-large and RoBERTa-large baseline detectors. This temporal module achieves the highest single-modality performance on the development set. Furthermore, a convex-weighted late fusion strategy improves overall performance to 0.804 and 0.669 macro-F1 on the development and test sets, respectively. The learned fusion effectively assigns zero weight to acoustics, demonstrating that conversational timing serves as a lightweight, interpretable complement for dyadic depression screening.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。