提出可信评估框架,让新生儿癫痫检测AI性能更真实可靠。
Honest and Reliable Evaluation and Expert Equivalence Testing of Automated Neonatal Seizure Detection
- 用真实与合成标注数据测试多种指标,发现马修斯和皮尔逊相关系数在类别不平衡下表现更好
- 多评者图灵测试结合弗莱斯k值能准确反映AI达到专家水平的表现
- 建议报告平衡指标、敏感性特异性及多评者图灵测试结果,提升可比性
新生儿癫痫检测的机器学习模型评估必须可靠才能推动临床应用。现有方法常依赖不一致且有偏的指标,影响模型可比性和解释性。专家声称的AI性能往往缺乏严格验证,可靠性存疑。本研究系统评估了常见性能指标、共识策略与人类专家等效性测试,在不同类别不平衡、评者间一致性及评者数量条件下进行分析。结果表明,马修斯相关系数和皮尔逊相关系数在类别不平衡情况下优于曲线下面积(AUC)。共识类型对评者数量和一致性敏感。在多种人类专家等效性测试中,采用弗莱斯k值的多评者图灵测试最能捕捉达到专家水平的AI表现。建议报告:(1) 至少一个平衡指标,(2) 敏感性、特异性、阳性预测值和阴性预测值,(3) 基于弗莱斯k的多评者图灵测试结果,(4) 以上全部在独立验证集上。该框架为临床验证提供了重要前提,支持对新生儿癫痫检测AI方法进行全面而诚实的评估。
原文摘要 · Abstract (English)
Reliable evaluation of machine learning models for neonatal seizure detection is critical for clinical adoption. Current practices often rely on inconsistent and biased metrics, hindering model comparability and interpretability. Expert-level claims about AI performance are frequently made without rigorous validation, raising concerns about their reliability. This study aims to systematically evaluate common performance metrics and propose best practices tailored to the specific challenges of neonatal seizure detection. Using real and synthetic seizure annotations, we assessed standard performance metrics, consensus strategies, and human-expert level equivalence tests under varying class imbalance, inter-rater agreement, and number of raters. Matthews and Pearson's correlation coefficients outperformed the area under the curve in reflecting performance under class imbalance. Consensus types are sensitive to the number of raters and agreement level among them. Among human-expert level equivalence tests, the multi-rater Turing test using Fleiss k best captured expert-level AI performance. We recommend reporting: (1) at least one balanced metric, (2) Sensitivity, specificity, PPV and NPV, (3) Multi-rater Turing test results using Fleiss k, and (4) All the above on held-out validation set. This proposed framework provides an important prerequisite to clinical validation by enabling a thorough and honest appraisal of AI methods for neonatal seizure detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。