arXiv:2509.08188cs.LGcs.NE2025-09

对比生成模型合成脑电伪迹,发现对抗训练效果更好。

ArtifactGen: Benchmarking WGAN-GP vs Diffusion for Label-Aware EEG Artifact Synthesis

  • 用WGAN-GP和扩散模型生成带标签的脑电伪迹
  • WGAN-GP在频谱匹配上更接近真实数据,MMD更低
  • 两类模型分类恢复能力弱,适合研究条件增强机制

脑电图(EEG)中的伪迹(如肌肉、眼动、电极、咀嚼、颤抖)干扰自动化分析,但大规模标注成本高。本文研究现代生成模型能否合成真实且带标签的伪迹片段,用于数据增强与系统测试。基于TUH EEG伪迹(TUAR)数据集,构建按受试者划分的固定长度多通道窗口(如250样本),并为不同模型定制预处理:对抗训练使用每窗最小最大归一化,扩散模型采用每记录/通道的z-score标准化。对比带有投影判别器的条件WGAN-GP与带无分类器引导的1D去噪扩散模型,从三方面评估:(i) 保真度,通过韦尔奇频段功率差(Δδ, Δθ, Δα, Δβ)、通道协方差弗罗贝尼乌斯距离、自相关L2及分布指标(MMD/PRD);(ii) 特异性,通过轻量kNN/分类器进行类别条件恢复;(iii) 实用性,通过数据增强对伪迹识别的影响。结果显示,WGAN-GP在频谱对齐和降低MMD方面表现更优,但两类模型均存在弱类别条件恢复能力,限制即刻增强效果,揭示更强条件控制与覆盖的改进空间。我们发布可复现的流程——数据清单、训练配置与评估脚本——以建立脑电伪迹合成基准,并暴露未来工作的可行动失败模式。

原文摘要 · Abstract (English)

Artifacts in electroencephalography (EEG) -- muscle, eye movement, electrode, chewing, and shiver -- confound automated analysis yet are costly to label at scale. We study whether modern generative models can synthesize realistic, label-aware artifact segments suitable for augmentation and stress-testing. Using the TUH EEG Artifact (TUAR) corpus, we curate subject-wise splits and fixed-length multi-channel windows (e.g., 250 samples) with preprocessing tailored to each model (per-window min-max for adversarial training; per-recording/channel $z$-score for diffusion). We compare a conditional WGAN-GP with a projection discriminator to a 1D denoising diffusion model with classifier-free guidance, and evaluate along three axes: (i) fidelity via Welch band-power deltas ($Δδ,\ Δθ,\ Δα,\ Δβ$), channel-covariance Frobenius distance, autocorrelation $L_2$, and distributional metrics (MMD/PRD); (ii) specificity via class-conditional recovery with lightweight $k$NN/classifiers; and (iii) utility via augmentation effects on artifact recognition. In our setting, WGAN-GP achieves closer spectral alignment and lower MMD to real data, while both models exhibit weak class-conditional recovery, limiting immediate augmentation gains and revealing opportunities for stronger conditioning and coverage. We release a reproducible pipeline -- data manifests, training configurations, and evaluation scripts -- to establish a baseline for EEG artifact synthesis and to surface actionable failure modes for future work.

脑电伪迹生成模型数据增强对抗训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。