arXiv:2607.22980cs.LG2026-07

比较贝叶斯与频率派方法在跨被试脑电分类中的表现,发现贝叶斯完全池化提升有限。

Bayesian Complete-Pooling in Cross-Subject Classification for Motor Imagery Electroencephalogram

论文配图:Bayesian Complete-Pooling in Cross-Subject Classification for Motor Imagery Electroencephalogram
图 1 · 摘自论文原文
  • 用贝叶斯完全池化建模跨被试运动想象脑电信号
  • 贝叶斯模型可靠性略优但预测不确定性更高,整体性能无显著差异
  • 适合关注概率校准与模型可信度的研究者

脑机接口长期追求免校准运行,但现有分类器多仅关注判别能力,忽视预测概率的校准性——尤其在非平稳脑电信号下,分布偏移时点估计模型可能过度自信。本研究在20个数据集上对比贝叶斯完全池化模型与频率派基线,在左右手运动想象脑电分类任务中表现。六个频率派流程均配对一个具有相同特征工程的贝叶斯流程,通过马尔可夫链蒙特卡洛后验采样训练。主评价指标为分解为可靠性和分辨率的布里尔分数(Brier score),辅以AUROC(判别力)和香农熵(尖锐度)。所有指标均采用随机效应元分析(REML,Knapp-Hartung调整)检验,并通过留一法影响分析验证。结果表明:贝叶斯完全池化在可靠性上具统计显著性提升,但预测不确定性增加(尖锐度下降);布里尔分数、分辨率与判别力无显著差异。各指标间异质性较低,但可靠性结果对留一法删除敏感。计算成本方面,贝叶斯流程耗能约为频率派的13倍,但仍低于常见家用电器水平。结论显示:贝叶斯完全池化对跨被试运动想象分类的实用价值有限,未来应探索跨被试与会话的部分池化策略。

原文摘要 · Abstract (English)

Brain-computer interfaces (BCIs) have long sought calibration-free operation, but classifiers are typically benchmarked by discrimination alone, blind to whether predicted probabilities are well calibrated - a meaningful gap given nonstationary electroencephalogram (EEG) signals and the risk of overconfident point-estimate classifiers under distribution shift. We conducted a large-scale study contrasting Bayesian complete-pooling models against frequentist baselines for cross-subject, left-hand versus right-hand motor imagery EEG classification across 20 datasets. Six frequentist pipelines were each paired with an analogous Bayesian pipeline sharing identical feature engineering, fit via Markov chain Monte Carlo posterior sampling. Our primary metric was the Brier score, decomposed into reliability and resolution, alongside AUROC for discrimination and Shannon entropy for sharpness. Each metric was analyzed via random-effects meta-analysis (REML, Knapp-Hartung adjustment), verified by leave-one-out influence analysis. Bayesian complete-pooling produced statistically but not practically significant improvements in reliability and increases in predictive uncertainty (lower sharpness); Brier score, resolution, and discrimination showed no significant differences. Between-study heterogeneity was low across all metrics, though the reliability result was sensitive to leave-one-out removal. We additionally profiled computational cost, finding that Bayesian pipelines consumed roughly thirteen times more energy than their frequentist counterparts, a cost that remains modest relative to common household appliances. These results suggest that Bayesian complete-pooling alone offers limited practical benefit for cross-subject motor imagery classification, and that partial-pooling across subjects and sessions is a more promising direction for future work.

脑机接口贝叶斯方法跨被试概率校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。