提出新框架,统一解码不同人脑的语音声调。
Towards Homogeneous Lexical Tone Decoding from Heterogeneous Intracranial Recordings
- 分离个体差异与共性特征,统一建模跨人脑声调信号
- 在407个普通话音节上实现更优解码性能
- 适合需要跨患者通用解码的脑机接口研究
近年来脑机接口技术已能从颅内记录中解码词汇声调,为失语的声调语言使用者恢复沟通能力带来可能。然而,生理与仪器因素导致的数据异质性,给统一的侵入式声调解码带来挑战。传统个体特异性模型在异质解码范式下无法捕捉通用神经表征,且难以跨被试利用数据。为此,我们提出同质-异质解耦学习框架(H2DiLR),从多被试颅内记录中分离并学习同质性与异质性。为验证该方法,我们采集了多名参与者阅读普通话材料的立体脑电图(sEEG)数据,涵盖407个音节,几乎覆盖所有普通话汉字。大量实验表明,作为统一解码范式,H2DiLR显著优于传统异质解码方法;同时实证证明其能有效捕获神经表征学习中的同质性与异质性。
原文摘要 · Abstract (English)
Recent advancements in brain-computer interfaces (BCIs) have enabled the decoding of lexical tones from intracranial recordings, offering the potential to restore the communication abilities of speech-impaired tonal language speakers. However, data heterogeneity induced by both physiological and instrumental factors poses a significant challenge for unified invasive brain tone decoding. Traditional subject-specific models, which operate under a heterogeneous decoding paradigm, fail to capture generalized neural representations and cannot effectively leverage data across subjects. To address these limitations, we introduce Homogeneity-Heterogeneity Disentangled Learning for neural Representations (H2DiLR), a novel framework that disentangles and learns both the homogeneity and heterogeneity from intracranial recordings across multiple subjects. To evaluate H2DiLR, we collected stereoelectroencephalography (sEEG) data from multiple participants reading Mandarin materials comprising 407 syllables, representing nearly all Mandarin characters. Extensive experiments demonstrate that H2DiLR, as a unified decoding paradigm, significantly outperforms the conventional heterogeneous decoding approach. Furthermore, we empirically confirm that H2DiLR effectively captures both homogeneity and heterogeneity during neural representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。