用梯度提升融合语音特征,精准识别婴儿哭声以辅助心理评估
Enhancing Infant Crying Detection with Gradient Boosting for Improved Emotional and Mental Health Diagnostics
- 结合Wav2Vec与传统声学特征进行多模态输入
- 在真实数据集上分类准确率显著优于现有方法
- 适合医疗健康与儿童心理研究者参考
婴儿哭声可作为生理和情绪状态的重要指标。本文提出一种综合方法,从音频数据中检测婴儿哭声。通过将Wav2Vec与传统音频特征融合,并采用梯度提升机(Gradient Boosting Machines)进行哭声分类。在真实世界数据集上验证该方法,结果表明其性能显著优于现有技术。
原文摘要 · Abstract (English)
Infant crying can serve as a crucial indicator of various physiological and emotional states. This paper introduces a comprehensive approach detecting infant cries within audio data. We integrate Wav2Vec with traditional audio features and employ Gradient Boosting Machines for cry classification. We validate our approach on a real world dataset, demonstrating significant performance improvements over existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。