arXiv:2605.00871eess.SPcs.AI2026-05中稿 · CVPR

动态调整时间尺度的医学信号模型,提升脑电分析精度与速度

NAKUL-Med: Spectral-Graph State Space Models with Dynamics Kernels for Medical Signals

论文配图:NAKUL-Med: Spectral-Graph State Space Models with Dynamics Kernels for Medical Signals
图 1 · 摘自论文原文
  • 通过可变核大小分支与元网络,自适应选择时间尺度
  • 在脑电运动想象任务中达91.7%准确率,参数少28%,推理快一倍
  • 融合频谱上下文与电极拓扑,适合脑电、心电等多通道医疗信号

状态空间模型(SSMs)虽具线性时间复杂度,但在多通道生理信号分析中存在三大局限:固定核无法捕捉多尺度时序动态(毫秒级准备期与数十毫秒执行瞬变),马尔可夫状态更新限制周期振荡的全局上下文建模,且通道独立处理忽略了电极空间拓扑。本文提出NAKUL,通过三方面改进:(1)动态核生成——并行不同核大小(3,5,7,11步)的SSM分支,由元网络根据输入统计加权,实现自适应时序尺度选择;(2)频谱上下文建模——基于FFT的可学习高斯频带滤波器,以$O(N \log N)$复杂度捕获全局周期模式;(3)图引导空间注意力——利用固定电极拓扑为多头注意力提供空间偏置,实现有原则的跨通道交互。在BCI竞赛IV-2a运动想象任务(主要基准)中,NAKUL达91.7±0.6%准确率,接近EEG-Conformer(92.1±0.7%),但参数减少28%(2.5M vs 3.5M),推理速度提升2.0×(4.3ms vs 8.7ms)。模型泛化至脑电情绪识别(83.6%)、多模态脑电-fMRI(91.4%)及医学影像(超声92.8%),展现架构通用性。消融实验显示动态核贡献+2.6%性能,且其选择模式与已知神经动力学相关联。

原文摘要 · Abstract (English)

State space models (SSMs) achieve linear-time complexity but struggle with multi-channel physiological signals due to three limitations: fixed kernels cannot capture multi-scale temporal dynamics (motor preparation over hundreds of milliseconds vs. execution transients in tens of milliseconds), Markovian state updates restrict global context for periodic oscillations, and channel-independent processing ignores spatial electrode topology. We introduce NAKUL, extending SSMs for medical signal analysis through three contributions: (1) Dynamic Kernel Generation-parallel SSM branches with varying kernel sizes (3, 5, 7, 11 timesteps) are weighted by a meta-network that analyzes input statistics, enabling adaptive temporal scale selection; (2) Spectral Context Modeling-FFT-based operations with learnable Gaussian frequency band filters capture global periodic patterns in $O(N \log N)$ complexity; (3) Graph-Guided Spatial Attention-fixed electrode topology provides spatial biases to multi-head attention for principled cross-channel interaction. On BCI Competition IV-2a motor imagery (our primary benchmark), NAKUL achieves 91.7$\pm$0.6\% accuracy, matching EEG-Conformer (92.1$\pm$0.7\%) while using 28\% fewer parameters (2.5M vs 3.5M) and 2.0$\times$ faster inference (4.3ms vs 8.7ms). The model generalizes to EEG emotion recognition (83.6\%), multimodal EEG-fMRI (91.4\%), and medical imaging (92.8\% on ultrasound), demonstrating architectural versatility. Ablations show dynamic kernels contribute +2.6\% and exhibit interpretable scale selection patterns correlated with known neural dynamics.

状态空间模型脑电分析动态核医疗信号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。