arXiv:2604.24096cs.LGcs.AI2026-04中稿 · ed

通过多样数据划分提升模型多样性,显著改善呼吸音分类效果。

Meta-Ensemble Learning with Diverse Data Splits for Improved Respiratory Sound Classification

论文配图:Meta-Ensemble Learning with Diverse Data Splits for Improved Respiratory Sound Classification
图 1 · 摘自论文原文
  • 用不同数据划分方式训练基模型,增强预测差异性。
  • 在ICBHI数据集上达66.49分,优于现有方法。
  • 适合需要高鲁棒性的医疗语音分类场景。

由于数据集规模有限且受试者多样性不足,训练可靠的呼吸音分类模型仍具挑战。集成方法虽可提升鲁棒性,但若基模型使用相同数据训练,则易过拟合并产生高度相关预测,削弱集成效果。本文提出一种元集成学习方法:通过在不同数据划分下训练基模型,并用可训练的元模型融合其输出,以增强预测多样性。具体在ICBHI数据集上,采用固定80-20%划分和五折交叉验证两种分割策略,结合患者级与样本级两种数据粒度设置。该策略使基模型预测更具差异性,从而提升元模型泛化能力。实验表明,该方法在ICBHI基准上取得新最佳性能(分数66.49%),并在两个分布外数据集上表现更优,展现出在真实临床数据中的应用潜力。

原文摘要 · Abstract (English)

Training reliable respiratory sound classification models remains challenging due to the limited size and subject diversity of datasets. Ensemble methods can improve robustness, but when base models are trained on identical data, models tend to overfit and produce highly correlated predictions, thereby reducing the effectiveness of ensembling. In this work, we investigate a meta-ensemble learning methodology that enhances prediction diversity by training base models on diverse data splits and combining their outputs through a trained meta-model. Specifically, we train base models on the ICBHI dataset using two data split settings: fixed 80-20% split and five-fold cross-validation split, under two data granularity settings: patient- and sample-level. The resulting diversity in base model predictions enables the meta-model to better generalize. Our approach achieves new state-of-the-art performance on the ICBHI benchmark, reaching a Score of 66.49% and showing improved generalization on two out-of-distribution datasets, indicating its potential applicability to real-world clinical data.

呼吸音分类集成学习医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。