arXiv:2501.01117cs.SDcs.AI2025-01被引 11

用深度神经决策树与森林模型,从咳嗽声中精准识别新冠,跨数据集表现稳定。

Robust COVID-19 Detection from Cough Sounds using Deep Neural Decision Tree and Forest: A Comprehensive Cross-Datasets Evaluation

  • 结合特征筛选与贝叶斯优化,构建深度神经决策树与森林模型。
  • 在5个数据集上均达0.92以上AUC,融合数据集达0.97。
  • 揭示地域与人群差异,支持多源数据融合提升泛化能力。

本研究提出一种鲁棒的新冠咳嗽声分类方法,采用深度神经决策树与深度神经决策森林,实现跨多种咳嗽声音数据集的稳定表现。通过全面提取音频特征,结合递归特征消除与交叉验证筛选关键特征,利用贝叶斯优化调优超参数,并在训练中引入SMOTE以平衡正负样本。通过阈值优化最大化ROC-AUC得分。在剑桥、Coswara、COUGHVID、Virufy及合并的Virufy与NoCoCoDa数据集上,分别取得0.97、0.98、0.92、0.93、0.99和0.99的显著AUC值。合并所有数据后,深度神经决策森林模型达到0.97 AUC。跨数据集分析揭示了新冠咳嗽声在人口与地理层面的差异,凸显特征迁移挑战,也表明数据集成可有效提升模型泛化性与检测能力。

原文摘要 · Abstract (English)

This research presents a robust approach to classifying COVID-19 cough sounds using cutting-edge machine-learning techniques. Leveraging deep neural decision trees and deep neural decision forests, our methodology demonstrates consistent performance across diverse cough sound datasets. We begin with a comprehensive extraction of features to capture a wide range of audio features from individuals, whether COVID-19 positive or negative. To determine the most important features, we use recursive feature elimination along with cross-validation. Bayesian optimization fine-tunes hyper-parameters of deep neural decision tree and deep neural decision forest models. Additionally, we integrate the SMOTE during training to ensure a balanced representation of positive and negative data. Model performance refinement is achieved through threshold optimization, maximizing the ROC-AUC score. Our approach undergoes a comprehensive evaluation in five datasets: Cambridge, Coswara, COUGHVID, Virufy, and the combined Virufy with the NoCoCoDa dataset. Consistently outperforming state-of-the-art methods, our proposed approach yields notable AUC scores of 0.97, 0.98, 0.92, 0.93, 0.99, and 0.99 across the respective datasets. Merging all datasets into a combined dataset, our method, using a deep neural decision forest classifier, achieves an AUC of 0.97. Also, our study includes a comprehensive cross-datasets analysis, revealing demographic and geographic differences in the cough sounds associated with COVID-19. These differences highlight the challenges in transferring learned features across diverse datasets and underscore the potential benefits of dataset integration, improving generalizability and enhancing COVID-19 detection from audio signals.

语音识别新冠检测深度学习数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。