arXiv:2508.19308cs.SD2025-08

轻量级哭声检测模型,噪声下仍保持高精度。

Infant Cry Detection In Noisy Environment Using Blueprint Separable Convolutions and Time-Frequency Recurrent Neural Network

  • 用分离卷积降低计算量,结合时频循环网络自适应去噪。
  • 在不同信噪比下准确率与F1均超越主流方法。
  • 适合嵌入式婴儿监护系统,尤其适用于嘈杂环境。

婴儿哭声检测是婴幼儿照护系统的关键环节。本文提出一种轻量且鲁棒的哭声检测方法,利用蓝图分离卷积减少计算复杂度,并采用时频循环神经网络实现自适应降噪。整体框架为多尺度卷积循环神经网络,通过高效的空域注意力机制和对比感知通道注意力模块,从对数梅尔谱图输入中提取局部与全局特征。使用多个公开数据集构建多样且具代表性的数据集,并通过环境污染技术生成真实场景中的噪声样本。实验表明,在不同信噪比条件下,该方法在准确率、F1分数和计算复杂度方面均优于多种先进方法。代码已开源:https://github.com/fhfjsd1/ICD_MMSP。

原文摘要 · Abstract (English)

Infant cry detection is a crucial component of baby care system. In this paper, we propose a lightweight and robust method for infant cry detection. The method leverages blueprint separable convolutions to reduce computational complexity, and a time-frequency recurrent neural network for adaptive denoising. The overall framework of the method is structured as a multi-scale convolutional recurrent neural network, which is enhanced by efficient spatial attention mechanism and contrast-aware channel attention module, and acquire local and global information from the input feature of log Mel-spectrogram. Multiple public datasets are adopted to create a diverse and representative dataset, and environmental corruption techniques are used to generate the noisy samples encountered in real-world scenarios. Results show that our method exceeds many state-of-the-art methods in accuracy, F1-score, and complexity under various signal-to-noise ratio conditions. The code is at https://github.com/fhfjsd1/ICD_MMSP.

音频处理轻量化婴儿监护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。