用相关性分析生成高保真合成脑电数据,助力心理疾病研究。
A Statistical Approach for Synthetic EEG Data Generation
- 基于频段间相关性构建结构,随机采样生成合成数据。
- 合成数据与真实数据相关系数相似,分类器区分率接近随机。
- 适合需要隐私保护的大规模脑电数据增强场景。
脑电图(EEG)数据对精神健康诊断至关重要,但大规模采集成本高、耗时长。合成数据生成为机器学习数据集扩充提供了可能,但保持情绪与心理健康信号的高质量合成仍具挑战。本研究提出结合相关性分析与随机采样生成真实感合成EEG的方法。首先通过相关性分析揭示各频率波段间的依赖关系,据此指导合成样本生成。保留与真实数据高度相关的样本,并通过分布分析与分类任务评估。使用随机森林模型区分合成与真实数据,其性能接近随机水平,表明合成数据具有高保真度。生成数据在统计与结构特性上与原始数据高度一致,相关系数相近,且PERMANOVA检验无显著差异。该方法提供了一种可扩展、隐私友好的EEG数据增强方案,有助于提升心理健康研究中的模型训练效率。
原文摘要 · Abstract (English)
Electroencephalogram (EEG) data is crucial for diagnosing mental health conditions but is costly and time-consuming to collect at scale. Synthetic data generation offers a promising solution to augment datasets for machine learning applications. However, generating high-quality synthetic EEG that preserves emotional and mental health signals remains challenging. This study proposes a method combining correlation analysis and random sampling to generate realistic synthetic EEG data. We first analyze interdependencies between EEG frequency bands using correlation analysis. Guided by this structure, we generate synthetic samples via random sampling. Samples with high correlation to real data are retained and evaluated through distribution analysis and classification tasks. A Random Forest model trained to distinguish synthetic from real EEG performs at chance level, indicating high fidelity. The generated synthetic data closely match the statistical and structural properties of the original EEG, with similar correlation coefficients and no significant differences in PERMANOVA tests. This method provides a scalable, privacy-preserving approach for augmenting EEG datasets, enabling more efficient model training in mental health research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。