arXiv:2606.24087cs.LG2026-06中稿 · MICCAI 2026被引 1

用确定性流匹配模型,从脑电波精准还原连续语音。

NeuroSonic: Conditional Flow Matching for EEG-to-Speech Reconstruction

论文配图:NeuroSonic: Conditional Flow Matching for EEG-to-Speech Reconstruction
图 1 · 摘自论文原文
  • 设计确定性流模型,直接推导脑电信号到语音的连续映射。
  • 在跨被试评估中,语音感知质量提升26.3%,尤其在噪声干扰段表现更优。
  • 适合脑机接口、神经解码领域研究者,可稳定生成高质量语音。

从头皮脑电图(EEG)重建连续语音仍面临根本挑战。EEG对皮层活动的测量具有弱信号、空间弥散和高度个体差异性,而语音则具有强谐波与时间结构的连贯声学轨迹。这种不匹配导致波形回归不稳定,且多步随机生成易受伪影和被试差异影响。本文提出NeuroSonic,一种基于条件流匹配的EEG-to-speech重建框架。不同于直接预测波形或通过随机去噪优化,NeuroSonic学习一个确定性的概率流速度场,在EEG条件下将噪声污染的声学状态逐步推向干净语音。EEG与音频被嵌入共享标记空间,并由时序条件门控Transformer参数化传输常微分方程。该方法显式建模轨迹演化,避免迭代随机采样。在CineBrain和EAV基准上进行跨被试评估,结果表明,相比代表性GAN、扩散模型与均值流基线,所提方法在分布真实性、频谱保真度与感知质量上均有提升,整体感知质量最高提升26.3%。性能差距在伪影密集段尤为明显,表明条件变化最剧烈时,确定性条件传输仍具鲁棒性。代码已开源。

原文摘要 · Abstract (English)

Reconstructing continuous speech from scalp electroencephalography (EEG) remains fundamentally challenging. EEG provides a weak, spatially diffuse, and highly variable measurement of distributed cortical activity, whereas speech is organized as a coherent acoustic trajectory with strong harmonic and temporal structure. The resulting mismatch makes waveform regression unstable and causes stochastic multi-step generation to be sensitive to artifact-dependent conditioning and subject variability. We introduce NeuroSonic, a conditional flow-matching framework for EEG-to-speech reconstruction. Instead of predicting waveforms directly or refining them through stochastic denoising, NeuroSonic learns a deterministic probability-flow velocity field that transports a noise-corrupted acoustic state toward clean speech under EEG conditioning. EEG and audio are embedded into a shared token space and processed by a time-conditioned gated Transformer that parameterizes the transport ordinary differential equation. This formulation models trajectory evolution explicitly while avoiding iterative stochastic sampling. We evaluate NeuroSonic on the CineBrain and EAV benchmarks under cross-subject evaluation. Across both datasets, the proposed method improves distributional realism, spectral fidelity, and perceptual quality over representative GAN-, diffusion-, and mean-flow baselines, with up to a 26.3\% gain in overall perceptual quality. The performance gap is most evident in artifact-heavy segments, where conditioning variability is strongest. These findings indicate that deterministic conditional transport provides a stable and effective formulation for EEG-driven speech reconstruction. Code is available at https://github.com/Y-Research-SBU/NeuroSonic/ .

脑机接口语音重建流匹配EEG解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。