arXiv:2603.12628q-bio.NCcs.AI2026-03

首次实现中文发音与听觉的统一脑转文本解码,支持跨模态对比。

Towards unified brain-to-text decoding across speech production and perception

  • 基于拼音组件分类+大模型后训练,实现单字数据上句级解码
  • 70亿参数模型性能超越数百亿参数商业模型,支持未见字符
  • 揭示发音比听觉激活更广脑区,且两模态神经模式相似但有时间差

言语生成与感知是人类日常交流的主要方式。以往脑转文本研究多聚焦单一模态及拼音文字。本文提出面向汉语普通话的统一脑转句子解码框架,可同时处理言语生成与感知。该框架具备强泛化能力,在仅用单字数据训练时即可实现句级解码,并支持训练中未出现的字符与音节。解码流程先从神经信号中分类声母与韵母(按汉语拼音),再由微调过的70亿参数大语言模型将无调拼音序列映射为中文句子。通过三阶段后训练与两阶段推理设计,其整体性能超过参数量达数百亿的商用大模型。观察到:言语生成激活更广泛的皮层区域;双模态共响应通道呈现相似活动模式,但感知存在相对生成的时间延迟;两半球解码性能基本相当。本工作不仅验证了统一解码框架的可行性,也为汉字语音的神经机制提供了新见解,推动多模态神经语言解码系统发展。

原文摘要 · Abstract (English)

Speech production and perception are the main ways humans communicate daily. Prior brain-to-text decoding studies have largely focused on a single modality and alphabetic languages. Here, we present a unified brain-to-sentence decoding framework for both speech production and perception in Mandarin Chinese. The framework exhibits strong generalization ability, enabling sentence-level decoding when trained only on single-character data and supporting characters and syllables unseen during training. In addition, it allows direct and controlled comparison of neural dynamics across modalities. Mandarin speech is decoded by first classifying syllable components in Hanyu Pinyin, namely initials and finals, from neural signals, followed by a post-trained large language model (LLM) that maps sequences of toneless Pinyin syllables to Chinese sentences. To enhance LLM decoding, we designed a three-stage post-training and two-stage inference framework based on a 7-billion-parameter LLM, achieving overall performance that exceeds larger commercial LLMs with hundreds of billions of parameters or more. In addition, several characteristics were observed in Mandarin speech production and perception: speech production involved neural responses across broader cortical regions than auditory perception; channels responsive to both modalities exhibited similar activity patterns, with speech perception showing a temporal delay relative to production; and decoding performance was broadly comparable across hemispheres. Our work not only establishes the feasibility of a unified decoding framework but also provides insights into the neural characteristics of Mandarin speech production and perception. These advances contribute to brain-to-text decoding in logosyllabic languages and pave the way toward neural language decoding systems supporting multiple modalities.

脑机接口语音生成中文解码多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。