提出新基准DEPOOL,系统评估抑郁检测中语音聚合方法的稳定性。
Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study
- 构建6种聚合架构+6个预训练模型的72种组合测试
- 发现三分之一配置崩溃为单一预测结果,与模型层选择有关
- 强调鲁棒性比单次准确率更重要,适合临床语音研究者
基于语音的抑郁检测将短音频片段特征压缩为说话人级判断,这一过程称为时间聚合,但其独立研究较少。现有基准通常固定单一自监督编码器和特定层,导致性能提升可能源于整体流程而非聚合方法本身。本文提出DEPOOL,一个受控基准,对比六种聚合架构与六个冻结的语音骨干网络,在英语和汉语抑郁语料库上进行测试,每个配置自动学习关键骨干层而非手动指定。在72种配置中,三分之一配置表现为对所有说话人预测同一类别,该失败现象既与骨干模型相关也与聚合方法有关;且在单种子运行中表现稳定的架构,在多种子重复训练下变得不可靠。因此,时间聚合的基准评估应以对骨干模型和随机种子的鲁棒性为核心指标,而非仅关注单一流程的平均准确率。
原文摘要 · Abstract (English)
Speech-based depression detection compresses features from short audio segments into one speaker-level decision, a step called temporal aggregation rarely studied on its own. Most benchmarks fix a single self-supervised encoder and a single hand-picked layer, so a reported gain may reflect the pipeline rather than the aggregation method itself. We introduce DEPOOL, a controlled benchmark that compares six aggregation architectures with six frozen speech backbones on an English and a Mandarin depression corpus, where each configuration learns which backbone layers matter rather than fixing one by hand. Across the resulting 72-configuration grid, a third of configurations collapse into predicting a single class for every speaker, a failure tied to the backbone as much as to the method, and the architecture that is most stable in a single-seed run becomes unreliable when training repeats across seeds. Robustness to backbone and seed, rather than average accuracy across a single pipeline, should be a first-class benchmarking criterion for temporal aggregation in clinical speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。