arXiv:2501.00064cs.SDcs.LG2025-01被引 3

用声音混叠增强模型泛化能力,解决不同数据集间诊断不准的问题。

Lungmix: A Mixup-Based Strategy for Generalization in Respiratory Sound Classification

  • 通过音量和随机掩码混合声波,按语义插值标签生成新样本
  • 在3个数据集上使4分类准确率最高提升3.55%
  • 适合需要跨数据集部署的呼吸疾病智能诊断场景

呼吸音分类在呼吸道疾病诊断中至关重要。尽管深度学习模型在多个呼吸音数据集上表现良好,但我们的实验表明,一个数据集训练的模型往往难以泛化到其他数据集,主要原因是数据采集与标注存在不一致性。为此,我们提出一种名为Lungmix的新数据增强技术,受Mixup启发。Lungmix通过使用音量和随机掩码混合波形,并根据标签语义进行插值,生成增强数据,帮助模型学习更通用的表征。在ICBHI、SPR和HF三个数据集上的综合评估表明,Lungmix显著提升了模型对未见数据的泛化能力。特别是,在4类分类任务中,性能最高提升3.55%,达到与直接在目标数据集训练模型相当的水平。

原文摘要 · Abstract (English)

Respiratory sound classification plays a pivotal role in diagnosing respiratory diseases. While deep learning models have shown success with various respiratory sound datasets, our experiments indicate that models trained on one dataset often fail to generalize effectively to others, mainly due to data collection and annotation \emph{inconsistencies}. To address this limitation, we introduce \emph{Lungmix}, a novel data augmentation technique inspired by Mixup. Lungmix generates augmented data by blending waveforms using loudness and random masks while interpolating labels based on their semantic meaning, helping the model learn more generalized representations. Comprehensive evaluations across three datasets, namely ICBHI, SPR, and HF, demonstrate that Lungmix significantly enhances model generalization to unseen data. In particular, Lungmix boosts the 4-class classification score by up to 3.55\%, achieving performance comparable to models trained directly on the target dataset.

呼吸音分类数据增强泛化能力Mixup

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。