arXiv:2505.18191eess.SPcs.AI2025-05被引 4

28种癫痫检测模型大比拼,发现性能与实际表现差距巨大。

Quantifying the Generalization Gap in Seizure Detection: A Large-Scale Empirical Benchmark via the SzCORE Challenge

  • 用65名患者4360小时脑电数据,严格测试28种先进算法
  • 最优模型F1仅32%,灵敏度37%、精确率29%,泛化能力差
  • 揭示模型自报战绩与真实表现严重不符,亟需统一评测标准

可靠的长时程脑电图(EEG)自动癫痫检测仍是未解难题,现有模型常在跨患者或临床场景中失效。目前文献报告的高效率常无法在真实部署中复现。为系统评估这一泛化差距,我们开展大规模实证研究,评估28种前沿算法,涵盖传统特征工程与现代深度学习方法。通过组织竞赛收集算法,并使用包含65名受试者共4,360小时连续脑电记录的严格保留私有数据集进行评估。专家神经生理学家对这些记录进行标注,确立癫痫事件的真实标签。采用SzCORE框架中的事件级指标(包括敏感性、精确率、F1分数和每日假阳性率)进行评估。结果显示,现有先进方法表现差异显著,最高F1分数仅为32%(敏感性37%,精确率29%),凸显该任务持续存在的挑战。分析发现,峰值性能与群体稳定性之间存在不一致:取得最高总体F1的算法并未在各受试者间保持最稳定的排名。此次独立评估揭示了自我报告效能与保留集表现之间的明显差距,强调了建立标准化、严格基准评测体系的迫切需求。评估基础设施已转型为持续开放的基准平台,推动可复现研究,加速稳健癫痫检测算法的发展。

原文摘要 · Abstract (English)

Reliable automatic seizure detection from long-term electroencephalography (EEG) remains an unsolved challenge, as current models often fail to generalize across patients or clinical settings. Manual EEG review still is the standard of care, highlighting the need for robust models and standardized evaluation. The current literature often reports high efficacy, yet these models frequently fail when deployed to unseen patient populations. To rigorously assess this generalization gap, we conducted a large-scale empirical study evaluating 28 state-of-the-art algorithmic architectures, ranging from classical feature engineering to modern Deep Learning. These algorithms were collected by organizing a competition. A strictly held-out private dataset of continuous EEG recordings from 65 subjects, totaling 4,360 hours of data, was utilized to evaluate algorithm performance. Expert neurophysiologists annotated these recordings, establishing the ground truth for seizure events. Algorithms were evaluated using event-based metrics from the SzCORE framework, including sensitivity, precision, F1-score, and false positive rate per day. Results revealed significant performance variability among state-of-the-art approaches, with the top F1 score of 32% (sensitivity 37%, precision 29%), highlighting the persistent difficulty of this task. Analysis uncovered a discordance between peak performance and population-level stability. The algorithms achieving the highest aggregate F1-scores did not achieve the most consistent ranking across subjects. This independent evaluation exposed a notable gap between self-reported efficacies and hold-out performance, underscoring the critical need for standardized, rigorous benchmarking. The evaluation infrastructure transitions into a continuously open benchmarking platform, fostering reproducible research and accelerating robust seizure detection algorithm development.

癫痫检测脑电图泛化能力基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。