随机初始化的Transformer能有效平滑睡眠分期,无需训练即可超越传统方法。
Rethinking Random Transformers as Adaptive Sequence Smoothers for Sleep Staging

- 用随机注意力机制实现自适应序列平滑,兼顾全局平均与内容相似性。
- 未训练的Transformer在睡眠分期上表现优于启发式平滑,准确率显著提升。
- 适合追求高效、轻量级医疗边缘部署的睡眠分析系统开发者。
自动睡眠分期通常依赖Transformer模型,假设其能捕捉复杂的长程依赖关系。本文挑战这一观点,指出睡眠序列具有强局部时间连续性。研究发现,未经训练的随机初始化Transformer在睡眠分期任务中表现优异,且持续优于启发式平滑方法。通过提出随机注意力先验核(RAPK),我们证明随机自注意力机制可作为自适应平滑器,在保留阶段转换的同时平衡全局平均与内容相似性。基于局部平滑影响指数(LSII)和加权转换熵(WTE)两个指标,我们证实大多数基于Transformer的睡眠分期性能提升源于架构归纳偏置,而非参数学习。结果表明,睡眠分期可通过结构驱动的平滑机制有效解决,而非复杂依赖建模,从而推动更高效、适合边缘部署的大规模生理监测系统发展。
原文摘要 · Abstract (English)
Automatic sleep staging commonly adopts Transformers under the assumption that they learn complex long-range dependencies. We challenge this view by revealing a neglected property of sleep sequences: strong local temporal continuity. We show that a randomly initialized Transformer, without any training, substantially improves sleep staging performance and consistently outperforms heuristic smoothing. We formalize this effect via a Random Attention Prior Kernel (RAPK), showing that random self-attention acts as an adaptive smoother by balancing global averaging and content-based similarity while preserving stage transitions. Using two metrics, the Local Smoothness Influence Index (LSII) and the Weighted Transition Entropy (WTE), we provide evidence that most performance gains in Transformer-based sleep staging arise from architectural inductive bias rather than parameter learning. Our results suggest that sleep staging can be effectively addressed with structure-driven smoothing mechanisms rather than complex dependency modeling, enabling more efficient and edge-deployable healthcare systems for large-scale physiological monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。