用预训练模型提升肠道音模式识别准确率,助力肠胃病智能诊断
Benchmarking machine learning for bowel sound pattern classification from tabular features to pretrained models
- 对比表格特征、频谱图CNN和音频预训练模型三类方法
- 预训练模型在少样本类别上表现最优,区分肠音与非肠音AUC达0.89
- 适合医学人工智能、可穿戴健康监测及消化系统疾病研究者
电子听诊器和可穿戴录音传感器的发展推动了肠道音(BS)信号的自动化分析,实现基于数据的肠音模式及其与病理关联的研究。本研究基于16名健康受试者采集的肠音数据集,该数据集按四种既定肠音模式进行标注,用于评估机器学习模型对肠音模式的检测与分类性能。模型涵盖使用表格特征、基于频谱图的卷积神经网络以及在大型音频数据集上预训练的模型。结果表明,预训练模型明显更优,尤其在小样本类别上表现突出:使用HuBERT模型区分肠音与非肠音的AUC达到0.89,使用Wav2Vec 2.0模型区分肠音模式的AUC同样为0.89。这些成果为深入理解肠音特征及未来基于机器学习的胃肠检查诊断应用奠定了基础。
原文摘要 · Abstract (English)
The development of electronic stethoscopes and wearable recording sensors opened the door to the automated analysis of bowel sound (BS) signals. This enables a data-driven analysis of bowel sound patterns, their interrelations, and their correlation to different pathologies. This work leverages a BS dataset collected from 16 healthy subjects that was annotated according to four established BS patterns. This dataset is used to evaluate the performance of machine learning models to detect and/or classify BS patterns. The selection of considered models covers models using tabular features, convolutional neural networks based on spectrograms and models pre-trained on large audio datasets. The results highlight the clear superiority of pre-trained models, particularly in detecting classes with few samples, achieving an AUC of 0.89 in distinguishing BS from non-BS using a HuBERT model and an AUC of 0.89 in differentiating bowel sound patterns using a Wav2Vec 2.0 model. These results pave the way for an improved understanding of bowel sounds in general and future machine-learning-driven diagnostic applications for gastrointestinal examinations
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。