arXiv:2603.18299cs.LGcs.NE2026-03

用对抗学习让脑机接口跨会话稳定识音

ALIGN: Adversarial Learning for Generalizable Speech Neuroprosthesis

  • 通过对抗训练让编码器忽略会话特有干扰信号
  • 在未见会话上降低音素错误率与词错误率
  • 适合长期部署的脑机接口系统开发者

皮层内脑-机接口(BCI)在跨会话数据联合训练下可实现高精度语音解码。但在实际部署中,模型需在无标签新会话中泛化,性能常因跨会话非平稳性(如电极位移、神经元更新、用户策略变化)下降。本文提出ALIGN框架,基于多域对抗神经网络的半监督跨会话自适应方法。ALIGN联合训练特征编码器、音素分类器与域分类器,通过对抗优化使编码器保留任务相关特征,抑制会话特异性信息。在皮层内语音解码任务上评估表明,相比基线方法,ALIGN在未见过的会话中泛化能力更强,显著降低音素错误率与词错误率。结果表明,对抗域对齐是缓解会话级分布偏移、实现鲁棒纵向BCI解码的有效途径。

原文摘要 · Abstract (English)

Intracortical brain-computer interfaces (BCIs) can decode speech from neural activity with high accuracy when trained on data pooled across recording sessions. In realistic deployment, however, models must generalize to new sessions without labeled data, and performance often degrades due to cross-session nonstationarities (e.g., electrode shifts, neural turnover, and changes in user strategy). In this paper, we propose ALIGN, a session-invariant learning framework based on multi-domain adversarial neural networks for semi-supervised cross-session adaptation. ALIGN trains a feature encoder jointly with a phoneme classifier and a domain classifier operating on the latent representation. Through adversarial optimization, the encoder is encouraged to preserve task-relevant information while suppressing session-specific cues. We evaluate ALIGN on intracortical speech decoding and find that it generalizes consistently better to previously unseen sessions, improving both phoneme error rate and word error rate relative to baselines. These results indicate that adversarial domain alignment is an effective approach for mitigating session-level distribution shift and enabling robust longitudinal BCI decoding.

脑机接口语音解码对抗学习跨会话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。