用谐波分析提升小数据下打鼾检测准确率
Improving snore detection under limited dataset through harmonic/percussive source separation and convolutional neural networks
- 通过谐波/打击声分离提取打鼾的谐波特征
- 小数据训练时性能优于传统特征,提升显著
- 适合医疗听觉诊断、数据稀缺场景使用
打鼾是阻塞性睡眠呼吸暂停综合征(OSAS)常见的声学生物标志物,具有重要的临床诊断与监测价值。无论打鼾类型如何,多数打鼾声均表现出可识别的谐波模式,表现为时间维度上的能量分布特征。本文提出一种新方法,通过谐波/打击声源分离(HPSS)分析输入声音的谐波内容,以区分单声道打鼾与非打鼾声。基于HPSS生成的谐波频谱作为输入,用于传统神经网络架构,旨在提升在小样本学习框架下的打鼾检测性能。实验对比两种场景:1)使用完整打鼾与干扰声数据集;2)仅使用约1%的数据量进行训练。在完整数据下,该方法表现与文献中其他特征相当;但在小数据场景中,所提谐波特征显著优于经典输入特征,所有测试架构性能均明显提升。结果表明,引入谐波内容能更可靠地学习到多数打鼾声普遍存在的时频特征,尤其在数据有限条件下优势突出。
原文摘要 · Abstract (English)
Snoring, an acoustic biomarker commonly observed in individuals with Obstructive Sleep Apnoea Syndrome (OSAS), holds significant potential for diagnosing and monitoring this recognized clinical disorder. Irrespective of snoring types, most snoring instances exhibit identifiable harmonic patterns manifested through distinctive energy distributions over time. In this work, we propose a novel method to differentiate monaural snoring from non-snoring sounds by analyzing the harmonic content of the input sound using harmonic/percussive sound source separation (HPSS). The resulting feature, based on the harmonic spectrogram from HPSS, is employed as input data for conventional neural network architectures, aiming to enhance snoring detection performance even under a limited data learning framework. To evaluate the performance of our proposal, we studied two different scenarios: 1) using a large dataset of snoring and interfering sounds, and 2) using a reduced training set composed of around 1% of the data material. In the former scenario, the proposed HPSS-based feature provides competitive results compared to other input features from the literature. However, the key advantage of the proposed method lies in the superior performance of the harmonic spectrogram derived from HPSS in a limited data learning context. In this particular scenario, using the proposed harmonic feature significantly enhances the performance of all the studied architectures in comparison to the classical input features documented in the existing literature. This finding clearly demonstrates that incorporating harmonic content enables more reliable learning of the essential time-frequency characteristics that are prevalent in most snoring sounds, even in scenarios where the amount of training data is limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。