EEGUnity统一管理25个脑电数据集,提升大模型研究效率
EEGUnity: Open-Source Tool in Facilitating Unified EEG Datasets Towards Large-Scale EEG Model
- 模块化工具自动解析、清洗、统一多源脑电数据
- 在25个数据集上验证,实现高效批量处理与高质量输出
- 开源工具适合脑电研究者快速构建统一数据集
随着脑电图(EEG)数据集数量增多和大规模EEG模型的发展,管理多样化的EEG数据需求日益迫切。然而,EEG数据在内容、元数据和格式上的高度异构性,给多数据集整合与大规模研究带来挑战。为此,本文提出EEGUnity——一个开源工具,包含EEG解析、校正、批量处理及大语言模型增强模块。该工具可实现智能数据结构推断、数据清洗与格式统一,保障数据质量与一致性,为大规模脑电研究提供可靠基础。在25个不同来源的EEG数据集上评估表明,EEGUnity在解析与处理方面表现高效灵活。项目代码已公开于github.com/Baizhige/EEGUnity。
原文摘要 · Abstract (English)
The increasing number of dispersed EEG dataset publications and the advancement of large-scale Electroencephalogram (EEG) models have increased the demand for practical tools to manage diverse EEG datasets. However, the inherent complexity of EEG data, characterized by variability in content data, metadata, and data formats, poses challenges for integrating multiple datasets and conducting large-scale EEG model research. To tackle the challenges, this paper introduces EEGUnity, an open-source tool that incorporates modules of 'EEG Parser', 'Correction', 'Batch Processing', and 'Large Language Model Boost'. Leveraging the functionality of such modules, EEGUnity facilitates the efficient management of multiple EEG datasets, such as intelligent data structure inference, data cleaning, and data unification. In addition, the capabilities of EEGUnity ensure high data quality and consistency, providing a reliable foundation for large-scale EEG data research. EEGUnity is evaluated across 25 EEG datasets from different sources, offering several typical batch processing workflows. The results demonstrate the high performance and flexibility of EEGUnity in parsing and data processing. The project code is publicly available at github.com/Baizhige/EEGUnity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。