arXiv:2603.02245eess.AScs.LG2026-03

用LMU建模婴儿啼哭时序,提升跨域分类准确率

LMU-Based Sequential Learning and Posterior Ensemble Fusion for Cross-Domain Infant Cry Classification

  • 多分支CNN融合MFCC、STFT与基频轮廓,用增强LMU捕捉时序动态
  • 跨数据集测试下宏平均F1提升,且支持设备端实时运行
  • 熵门加权融合后验概率,有效保留领域特征并降低偏差

由于信号短、非平稳性强、跨个体和数据集存在显著域偏移,解析婴儿啼哭原因仍具挑战。本文提出一种紧凑声学框架:在多分支卷积神经网络编码器中融合梅尔频率倒谱系数(MFCC)、短时傅里叶变换(STFT)特征与基频(F0)轨迹,并采用增强型勒让德记忆单元(LMU)建模时序动态。相比LSTM,LMU具备更少的循环参数,实现稳定序列建模且利于部署。为提升跨数据集泛化能力,引入校准后验集成融合方法,通过熵门控加权机制,在保留领域特异性知识的同时缓解数据集偏差。在Baby2020与Baby Crying数据集上的实验表明,该方法在跨域评估下实现了更高的宏平均F1值,同时支持泄漏感知划分与设备端实时监测。

原文摘要 · Abstract (English)

Decoding infant cry causes remains challenging for healthcare monitoring due to short nonstationary signals, limited annotations, and strong domain shifts across infants and datasets. We propose a compact acoustic framework that fuses mel-frequency cepstral coefficients (MFCCs), short-time Fourier transform (STFT) features, and fundamental-frequency (F0) contours within a multi-branch convolutional neural network (CNN) encoder, and models temporal dynamics using an enhanced Legendre Memory Unit (LMU). Compared to LSTMs, the LMU backbone provides stable sequence modeling with substantially fewer recurrent parameters, supporting efficient deployment. To improve cross-dataset generalization, we introduce calibrated posterior ensemble fusion with entropy-gated weighting to preserve domain-specific expertise while mitigating dataset bias. Experiments on Baby2020 and Baby Crying demonstrate improved macro-F1 under cross-domain evaluation, along with leakage aware splits and real-time feasibility for on-device monitoring.

婴儿啼哭时序建模跨域学习轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。