arXiv:2503.06247cs.SDcs.AI2025-03中稿 · ICASSP 2025被引 2

构建婴儿啼哭标注数据集,提出因果时序表征方法提升检测性能。

Infant Cry Detection Using Causal Temporal Representation

  • 基于因果时序表征设计无监督聚类算法CRSTC,缓解数据稀缺问题。
  • 新数据集使有监督模型达到当前最优效果,显著提升分类性能。
  • 适合婴儿照护、语音事件检测方向的研究者与开发者参考。

本文针对声学事件检测中婴儿啼哭识别面临的挑战——即在其他声音和背景噪音干扰下缺乏精确标注数据——提出两项贡献。一是构建一个用于啼哭分割的标注数据集,使有监督模型实现最先进的性能;二是提出一种新型无监督方法:因果时序表示稀疏转移聚类(CRSTC),基于因果时序表征,更一般性地应对数据稀缺问题。通过整合检测到的啼哭片段,显著提升了下游婴儿啼哭分类的表现,展示了该方法在婴儿照护应用中的潜力。

原文摘要 · Abstract (English)

This paper addresses a major challenge in acoustic event detection, in particular infant cry detection in the presence of other sounds and background noises: the lack of precise annotated data. We present two contributions for supervised and unsupervised infant cry detection. The first is an annotated dataset for cry segmentation, which enables supervised models to achieve state-of-the-art performance. Additionally, we propose a novel unsupervised method, Causal Representation Spare Transition Clustering (CRSTC), based on causal temporal representation, which helps address the issue of data scarcity more generally. By integrating the detected cry segments, we significantly improve the performance of downstream infant cry classification, highlighting the potential of this approach for infant care applications.

婴儿啼哭无监督学习时序建模声学事件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。