用竞赛方式训练新生儿脑电图异常分级模型,提升临床决策支持能力。
Machine-learning competition to grade EEG background patterns in newborns with hypoxic-ischaemic encephalopathy
- 基于多中心数据开展机器学习竞赛,构建新生儿脑电图分类模型。
- 深度学习模型在验证集表现更稳定,但所有模型在新数据上性能均下降。
- 强调大规模多样数据对模型泛化的重要性,适合医疗AI研究者参考。
机器学习(ML)有望辅助专家监测高危新生儿的脑功能。然而,高质量标注数据稀缺限制了准确可靠的模型开发。通过举办机器学习竞赛,可提供专家标注的数据集,促进研究者间直接比较模型,并利用众包优势汇集多样化专业技能。本研究整合了来自多中心研究的102名新生儿的353小时脑电图(EEG)数据,已完成匿名化处理,并划分为训练、测试和独立保留的验证集。专家对脑电图背景模式的异常严重程度进行分级。随后,我们搭建了基于网络的竞赛平台,组织了一场针对新生儿脑电图背景模式严重程度分类的机器学习竞赛。竞赛结束后,前4名模型在独立保留的验证集上进行离线评估。尽管基于特征的模型在测试集上排名第一,但深度学习模型在验证集上泛化表现更好。所有方法在验证集上的性能均显著低于测试集表现,凸显模型在未见数据上泛化的挑战,强调在新生儿脑电图的机器学习研究中使用保留验证集的重要性。研究强调,必须在大规模且多样化的数据上训练模型,以确保稳健的泛化能力。竞赛结果展示了开放数据与协作式机器学习开发在推动新生儿神经监测临床决策支持工具发展方面的潜力。
原文摘要 · Abstract (English)
Machine learning (ML) has the potential to support and improve expert performance in monitoring the brain function of at-risk newborns. Developing accurate and reliable ML models depends on access to high-quality, annotated data, a resource in short supply. ML competitions address this need by providing researchers access to expertly annotated datasets, fostering shared learning through direct model comparisons, and leveraging the benefits of crowdsourcing diverse expertise. We compiled a retrospective dataset containing 353 hours of EEG from 102 individual newborns from a multi-centre study. The data was fully anonymised and divided into training, testing, and held-out validation datasets. EEGs were graded for the severity of abnormal background patterns. Next, we created a web-based competition platform and hosted a machine learning competition to develop ML models for classifying the severity of EEG background patterns in newborns. After the competition closed, the top 4 performing models were evaluated offline on a separate held-out validation dataset. Although a feature-based model ranked first on the testing dataset, deep learning models generalised better on the validation sets. All methods had a significant decline in validation performance compared to the testing performance. This highlights the challenges for model generalisation on unseen data, emphasising the need for held-out validation datasets in ML studies with neonatal EEG. The study underscores the importance of training ML models on large and diverse datasets to ensure robust generalisation. The competition's outcome demonstrates the potential for open-access data and collaborative ML development to foster a collaborative research environment and expedite the development of clinical decision-support tools for neonatal neuromonitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。