arXiv:2505.00839cs.SDcs.SI2025-05

构建听觉冥想数据集并用对比学习提升情绪与生理状态识别精度

SMSAT: A Multimodal Acoustic Dataset and Deep Contrastive Learning Framework for Affective and Physiological Modeling of Spiritual Meditation

  • 基于对比学习设计音频编码器,从声学时序数据提取高区分性特征
  • 在三种听觉条件下实现99.99%的情绪状态分类准确率
  • 适合关注冥想、心理健康与生物信号分析的研究者

理解听觉刺激如何影响情绪与生理状态,是推动情感计算与心理健康技术的关键。本文通过一套综合的生物信号测量方法,对三种听觉条件——精神冥想(SM)、音乐(M)和自然寂静(NS)——进行多模态评估。为此,我们提出了全新的SMSAT数据集,包含受控暴露协议下采集的声学时序(ATS)信号,注重人群多样性与实验一致性。为建模听觉诱导状态,我们开发了基于对比学习的SMSAT音频编码器,从ATS数据中提取高度判别性嵌入,在跨类与类内评估中达到99.99%的分类准确率。此外,提出融合25个手工与学习特征的沉静度分析模型(CAM),在不同听觉条件下实现稳定的99.99%分类准确率。相比现有最先进方法最高仅达90%的准确率,本模型性能显著提升。方差分析显示,冥想条件下的心率反应特征(CRC)差异更显著。该研究贡献了一个经过验证的多模态数据集与可扩展的深度学习框架,适用于压力监测、心理福祉与基于音频的干预应用。

原文摘要 · Abstract (English)

Understanding how auditory stimuli influence emotional and physiological states is fundamental to advancing affective computing and mental health technologies. In this paper, we present a multimodal evaluation of the affective and physiological impacts of three auditory conditions, that is, spiritual meditation (SM), music (M), and natural silence (NS), using a comprehensive suite of biometric signal measures. To facilitate this analysis, we introduce the Spiritual, Music, Silence Acoustic Time Series (SMSAT) dataset, a novel benchmark comprising acoustic time series (ATS) signals recorded under controlled exposure protocols, with careful attention to demographic diversity and experimental consistency. To model the auditory induced states, we develop a contrastive learning based SMSAT audio encoder that extracts highly discriminative embeddings from ATS data, achieving 99.99% classification accuracy in interclass and intraclass evaluations. Furthermore, we propose the Calmness Analysis Model (CAM), a deep learning framework integrating 25 handcrafted and learned features for affective state classification across auditory conditions, attaining robust 99.99% classification accuracy. In contrast, pairwise t tests reveal significant deviations in cardiac response characteristics (CRC) between SM analysis via ANOVA inducing more significant physiological fluctuations. Compared to existing state of the art methods reporting accuracies up to 90%, the proposed model demonstrates substantial performance gains (up to 99%). This work contributes a validated multimodal dataset and a scalable deep learning framework for affective computing applications in stress monitoring, mental well-being, and therapeutic audio-based interventions.

多模态情绪识别冥想研究对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。