arXiv:2607.18614cs.SDcs.HC2026-07

用马尔可夫模型建模听觉注意力序列,提升脑电解码准确率。

End-to-End Markov State Sequence Learning for Auditory Attention Decoding

论文配图:End-to-End Markov State Sequence Learning for Auditory Attention Decoding
图 1 · 摘自论文原文
  • 将注意力状态视为时序马尔可夫过程,联合优化神经发射与转移概率。
  • 在1秒窗口下达到86.5%(因果)和92.4%(非因果)准确率。
  • 适合需要连续注意力解码的脑机接口与智能助听设备研究者。

听觉注意力解码(AAD)从脑电图(EEG)等神经反应中识别听者关注的说话人,是神经控制助听器的关键技术。然而,现有模型多为独立短窗分类器,而听觉注意力具有时间持续性,且短窗的EEG-音频证据常噪声大、模糊。本文提出端到端马尔可夫AAD框架,基于条件随机场(CRF),在双状态注意力先验下训练窗口级神经发射。该框架将任意后端模型的输出对数作为马尔可夫发射,初始化转移率来自标准HMM,联合优化交叉熵与CRF目标,使时间连续性引导表征学习而非仅用于训练后平滑。我们还提出ESCNet,一种保留时间对齐特征的EEG-语音相关性主干,将两个均值皮尔逊相关系数差转换为状态对数。在动态AVGC数据集上,该方法优于事后HMM平滑;结合ESCNet,在1秒窗口下实现86.5%(因果)和92.4%(非因果)准确率。在静态KUL和USTC数据集上,相比固定率事后HMM基线,因果解码分别提升5.6%和2.0%,验证了将AAD建模为注意力序列优于孤立窗口分类。

原文摘要 · Abstract (English)

Auditory attention decoding (AAD) identifies the speaker a listener attends to from neural responses like electroencephalography (EEG), making it a key algorithm in neuro-steered hearing aids. However, most neural AAD models are trained as independent short-window classifiers, despite auditory attention being a temporally persistent cognitive state and short-window EEG--audio evidence often being noisy and ambiguous. We propose an end-to-end Markov AAD framework based on conditional random field (CRF) that trains window-level neural emissions under a two-state attention prior. The framework treats the logits of any AAD backbone as Markov emissions, learns the transition rate from a standard HMM initialization, and jointly optimizes cross-entropy and CRF objectives, allowing temporal continuity to guide representation learning rather than merely smoothing predictions after training. We also introduce ESCNet, an EEG--speech correlation backbone that preserves time-aligned features and converts the difference between two mean Pearson correlations into state logits. We evaluate the framework with four emission backbones spanning correlation-based, convolutional, recurrent, and attention-based designs. On the dynamic AVGC dataset, CRF training generally outperforms post-hoc HMM smoothing; with ESCNet, it achieves $86.5\%$ causal and $92.4\%$ non-causal accuracy using $1$s windows. On the static KUL and USTC datasets, it improves causal decoding over fixed-rate post-hoc HMM baselines by $5.6\%$ and $2.0\%$, respectively, showing the superiority of learning AAD as attention state sequence over isolated-window classification.

注意力解码脑机接口时序建模马尔可夫模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。