把语音中的压力看作动态变化过程,提升检测准确率。
Dynamic Stress Detection: A Study of Temporal Progression Modelling of Stress in Speech
- 从情绪标签推导细粒度压力标注,建模压力随时间演变
- 在MuSE上提升5%,在StressID上提升18%的准确率
- 适合关注心理状态动态监测的研究者与应用开发者
从语音中检测心理压力在高压场景中至关重要。以往研究多将压力视为静态标签,而本工作将压力建模为受历史情绪状态影响的时序演化现象。我们提出一种动态标注策略,从情绪标签推导出细粒度压力标注,并引入基于交叉注意力的序列模型——单向LSTM和Transformer编码器,以捕捉压力的时间演变。该方法在MuSE数据集上相比基线提升5%准确率,在StressID上提升18%,并在自建真实场景数据集上表现出良好泛化能力。结果表明,将压力视为动态构建物能显著提升检测性能。
原文摘要 · Abstract (English)
Detecting psychological stress from speech is critical in high-pressure settings. While prior work has leveraged acoustic features for stress detection, most treat stress as a static label. In this work, we model stress as a temporally evolving phenomenon influenced by historical emotional state. We propose a dynamic labelling strategy that derives fine-grained stress annotations from emotional labels and introduce cross-attention-based sequential models, a Unidirectional LSTM and a Transformer Encoder, to capture temporal stress progression. Our approach achieves notable accuracy gains on MuSE (+5%) and StressID (+18%) over existing baselines, and generalises well to a custom real-world dataset. These results highlight the value of modelling stress as a dynamic construct in speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。