arXiv:2604.05751eess.SPcs.LG2026-04被引 1

从脑电数据重建自然语音,关键在提取语调特征并用新模型增强表现。

Brain-to-Speech: Prosody Feature Engineering and Transformer-Based Reconstruction

  • 从颅内脑电信号中提取语调、音高、节奏等语调特征
  • 新提出的Transformer模型使语音可懂度和表达力显著提升
  • 适合神经假体与语言障碍康复研究者参考

本章提出一种新型脑到语音(BTS)合成方法,基于颅内脑电图(iEEG)数据,强调语调感知的特征工程与先进的Transformer模型以实现高保真语音重建。随着直接从脑活动解码语音的兴趣日益增长,本文融合神经科学、人工智能与信号处理,生成准确且自然的语音。我们设计了一种新流程,直接从复杂的脑电iEEG信号中提取关键语调特征,包括语调、音高与节奏。为有效利用这些重要特征生成自然语音,采用先进深度学习模型。此外,本文提出一种专为脑到语音任务设计的新型Transformer编码器架构。与传统模型不同,该架构整合提取的语调特征,显著提升语音重建质量,使生成语音具备更高的可懂度与表现力。详细评估显示,在定量与感知指标上均优于传统基线方法,如Griffin-Lim与基于CNN的重建。通过展示特征提取与Transformer学习的进展,本章推动了人工智能驱动的神经假体领域发展,为恢复言语障碍患者沟通能力的辅助技术铺平道路。最后,讨论了未来方向,包括扩散模型集成与实时推理系统。

原文摘要 · Abstract (English)

This chapter presents a novel approach to brain-to-speech (BTS) synthesis from intracranial electroencephalography (iEEG) data, emphasizing prosody-aware feature engineering and advanced transformer-based models for high-fidelity speech reconstruction. Driven by the increasing interest in decoding speech directly from brain activity, this work integrates neuroscience, artificial intelligence, and signal processing to generate accurate and natural speech. We introduce a novel pipeline for extracting key prosodic features directly from complex brain iEEG signals, including intonation, pitch, and rhythm. To effectively utilize these crucial features for natural-sounding speech, we employ advanced deep learning models. Furthermore, this chapter introduces a novel transformer encoder architecture specifically designed for brain-to-speech tasks. Unlike conventional models, our architecture integrates the extracted prosodic features to significantly enhance speech reconstruction, resulting in generated speech with improved intelligibility and expressiveness. A detailed evaluation demonstrates superior performance over established baseline methods, such as traditional Griffin-Lim and CNN-based reconstruction, across both quantitative and perceptual metrics. By demonstrating these advancements in feature extraction and transformer-based learning, this chapter contributes to the growing field of AI-driven neuroprosthetics, paving the way for assistive technologies that restore communication for individuals with speech impairments. Finally, we discuss promising future research directions, including the integration of diffusion models and real-time inference systems.

脑机接口语音合成Transformer神经假体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。