构建大规模颅内脑电数据集,助力癫痫定位与机器学习研究
Omni-iEEG: A Large-Scale, Comprehensive iEEG Dataset and Benchmark for Epilepsy Research
- 整合多中心数据,统一格式与标注标准
- 涵盖302例患者、178小时高清记录及超3.6万条病理事件标注
- 支持临床可解释的模型评估,适合癫痫研究与AI算法开发者
全球逾5000万人患癫痫,三分之一患者药物难治性发作,手术是最佳治愈手段。精准定位致痫区依赖颅内脑电(iEEG),但临床流程仍受限于人工审阅耗时。现有数据驱动方法多基于单中心数据,格式不一、元数据缺失、缺乏标准化基准和病理事件标注,阻碍可复现性与跨中心验证。我们通过系统整合公开来源的异构iEEG数据,构建了大型预手术iEEG资源Omni-iEEG,包含302名患者、178小时高分辨率记录,统一临床元数据如发作起始区、切除范围与手术结果,并经认证癫痫专家验证。此外,提供超过3.6万条专家标注的病理事件,支持可靠生物标志物研究。Omni-iEEG定义了基于临床先验的有意义任务,采用统一评估指标,实现模型在临床相关场景下的系统评估。除基准测试外,还展示了端到端建模长段iEEG的潜力,以及非神经生理领域预训练表示的迁移能力。这些贡献使Omni-iEEG成为可复现、可泛化且具临床转化价值的癫痫研究基础。项目页面及数据代码链接见omni-ieeg.github.io/omni-ieeg。
原文摘要 · Abstract (English)
Epilepsy affects over 50 million people worldwide, and one-third of patients suffer drug-resistant seizures where surgery offers the best chance of seizure freedom. Accurate localization of the epileptogenic zone (EZ) relies on intracranial EEG (iEEG). Clinical workflows, however, remain constrained by labor-intensive manual review. At the same time, existing data-driven approaches are typically developed on single-center datasets that are inconsistent in format and metadata, lack standardized benchmarks, and rarely release pathological event annotations, creating barriers to reproducibility, cross-center validation, and clinical relevance. With extensive efforts to reconcile heterogeneous iEEG formats, metadata, and recordings across publicly available sources, we present $\textbf{Omni-iEEG}$, a large-scale, pre-surgical iEEG resource comprising $\textbf{302 patients}$ and $\textbf{178 hours}$ of high-resolution recordings. The dataset includes harmonized clinical metadata such as seizure onset zones, resections, and surgical outcomes, all validated by board-certified epileptologists. In addition, Omni-iEEG provides over 36K expert-validated annotations of pathological events, enabling robust biomarker studies. Omni-iEEG serves as a bridge between machine learning and epilepsy research. It defines clinically meaningful tasks with unified evaluation metrics grounded in clinical priors, enabling systematic evaluation of models in clinically relevant settings. Beyond benchmarking, we demonstrate the potential of end-to-end modeling on long iEEG segments and highlight the transferability of representations pretrained on non-neurophysiological domains. Together, these contributions establish Omni-iEEG as a foundation for reproducible, generalizable, and clinically translatable epilepsy research. The project page with dataset and code links is available at omni-ieeg.github.io/omni-ieeg.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。