用因果注意力提升婴儿哭声分类鲁棒性,减少环境干扰影响。
Domain-Agnostic Causal-Aware Audio Transformer for Infant Cry Classification
- 引入因果注意力与层级结构,学习更可靠的声学特征。
- 在两个数据集上准确率提升2.6%,宏平均F1提升2.2点。
- 适合临床早筛场景,对未知录音环境有强泛化能力。
准确且可解释的婴儿哭声副语言分类对新生儿窘迫的早期发现和临床决策支持至关重要。然而,现有深度学习方法依赖相关性驱动的声学表征,易受噪声、虚假线索及不同录音环境下的领域偏移影响。本文提出DACH-TIC:一种面向婴儿哭声分类的无领域依赖因果感知分层音频变压器。该模型在统一框架中融合因果注意力、层级表示学习、多任务监督与对抗性领域泛化。采用结构化Transformer骨干网络,包含局部词元级与全局语义编码器,结合因果注意力掩码与受控扰动训练以逼近反事实声学变化。领域对抗目标促使环境不变表征学习,多任务学习联合优化哭声类型识别、痛苦强度估计与因果相关性预测。在Baby Chillanto与Donate-a-Cry数据集上评估,并使用ESC-50环境噪声叠加进行领域增强。实验结果表明,DACH-TIC优于当前最优基线(如HTS-AT与SE-ResNet Transformer),准确率提升2.6个百分点,宏平均F1提升2.2点,同时增强因果一致性。模型在未见声学环境中表现稳定,领域性能差距仅2.4个百分点,适用于真实世界新生儿声学监测系统。
原文摘要 · Abstract (English)
Accurate and interpretable classification of infant cry paralinguistics is essential for early detection of neonatal distress and clinical decision support. However, many existing deep learning methods rely on correlation-driven acoustic representations, which makes them vulnerable to noise, spurious cues, and domain shifts across recording environments. We propose DACH-TIC, a Domain-Agnostic Causal-Aware Hierarchical Audio Transformer for robust infant cry classification. The model integrates causal attention, hierarchical representation learning, multi-task supervision, and adversarial domain generalization within a unified framework. DACH-TIC employs a structured transformer backbone with local token-level and global semantic encoders, augmented by causal attention masking and controlled perturbation training to approximate counterfactual acoustic variations. A domain-adversarial objective promotes environment-invariant representations, while multi-task learning jointly optimizes cry type recognition, distress intensity estimation, and causal relevance prediction. The model is evaluated on the Baby Chillanto and Donate-a-Cry datasets, with ESC-50 environmental noise overlays for domain augmentation. Experimental results show that DACH-TIC outperforms state-of-the-art baselines, including HTS-AT and SE-ResNet Transformer, achieving improvements of 2.6 percent in accuracy and 2.2 points in macro-F1 score, alongside enhanced causal fidelity. The model generalizes effectively to unseen acoustic environments, with a domain performance gap of only 2.4 percent, demonstrating its suitability for real-world neonatal acoustic monitoring systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。