arXiv:2606.08594cs.LGeess.SP2026-06

小模型已够用,但常用指标不靠谱,降噪效果未必提升脑机解码性能。

How Much Capacity Does EEG Denoising Need? Ultra-Compact Networks reveal Benchmark Saturation and Metric-Utility Gap

论文配图:How Much Capacity Does EEG Denoising Need? Ultra-Compact Networks reveal Benchmark Saturation and Metric-Utility Gap
图 1 · 摘自论文原文
  • 用极简网络控制容量变量,只调通道宽度测试降噪能力极限
  • 3-6.5K参数后性能饱和,后续每10倍参数仅增0.015相关系数
  • 重建好≠分类准,传统指标无法预测脑机接口实际表现

深度学习的脑电降噪模型参数量从数万增至数百万,但此前研究未将模型容量作为独立变量测试,也未验证重建指标是否反映下游神经信号实用性。本研究在固定架构、损失函数、数据划分和训练流程的前提下,仅调整通道宽度(1.05K至40.26K参数),使用最小深度可分离卷积U-Net进行测试。模型在EEGDenoiseNet基准、跨数据集脑机接口迁移测试、受控基线重训及九名BCI Competition IV-2a受试者的运动想象分类任务中评估,采用五种解码器。重建性能在3-6.5K参数时饱和,此后每增加一个对数单位参数,相关系数最多提升0.015。846万参数基线模型在相同训练流程下与40.26K参数紧凑模型在眼动伪迹去除上表现一致——200倍参数差距无优势;而补丁注意力控制组也呈现相似递减收益。下游评估揭示分类器依赖的指标-效用差距:以重建优化的降噪显著降低CSP+LDA分类性能,所有九名受试者三种伪迹类型下准确率下降(最佳去噪准确率0.547,原始噪声基线0.612;Bonferroni p=0.0488),自然记录试验中仍持续下降(Δ=-0.047;BH-FDR q=0.0049)。端到端神经解码器则表现各异或中性。标准脑电降噪基准远低于当前模型容量已趋饱和,且重建指标无法预测脑机接口实用性。超紧凑模型体积仅33-46 KB,每片段计算量1.27-2.61M FLOPs,适合边缘部署。研究呼吁开展容量可控评估、设计更难的任务导向基准,并强制进行下游验证。

原文摘要 · Abstract (English)

Deep learning EEG denoising architectures have scaled from tens of thousands to tens of millions of parameters, yet no prior study has isolated model capacity as the experimental variable or tested whether reconstruction metrics predict downstream neural-signal utility. We address both gaps by fixing architecture, loss, data split, and training recipe while sweeping only channel width from 1.05K to 40.26K parameters in a minimal depthwise-separable convolutional U-Net. Models were evaluated on the EEGDenoiseNet benchmark, cross-dataset BCI transfer tests, controlled baseline retraining, and downstream motor-imagery classification with five decoder families across all nine BCI Competition IV-2a subjects. Reconstruction performance saturated by 3-6.5K parameters, with post-elbow gains of at most 0.015 correlation coefficient per log10-parameter unit. An 8.46M-parameter baseline retrained under the same pipeline matched the 40.26K compact variant on EOG--a 200x parameter gap yielding no advantage--while a Patch-Transformer control reproduced the same diminishing-return shape. Downstream evaluation exposed a classifier-dependent metric-utility gap: reconstruction-optimized denoising significantly degraded CSP+LDA classification across all nine subjects and three artifact types (best denoised accuracy 0.547 vs. 0.612 noisy baseline; Bonferroni p=0.0488), persisting on naturally recorded trials (Delta=-0.047; BH-FDR q=0.0049). End-to-end neural decoders showed variable or neutral effects. Standard EEG denoising benchmarks are saturated far below current model capacity, and reconstruction metrics do not predict BCI utility. Ultra-compact models at 33-46 KB and 1.27-2.61M FLOPs/segment are practical for edge deployment. These findings argue for capacity-controlled evaluation, harder task-aware benchmarks, and mandatory downstream validation.

脑电降噪模型容量指标有效性边缘部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。