用扩散模型从脑信号解码口型,实现动态说话脸重建。
Towards Dynamic Neural Communication and Speech Neuroprosthesis Based on Viseme Decoding
- 基于扩散模型,从非侵入式脑信号中解码口型动作。
- 在连续语句上成功重建连贯的唇部运动轨迹。
- 为瘫痪患者提供动态语音神经假体,适合脑机接口研究者。
从人类神经信号中解码文字、语音或图像,对患者神经假体和普通用户创新通信工具具有广阔前景。尽管神经信号包含语音意图、动作和音素细节等多重信息,但生成有意义输出仍具挑战,现有研究多聚焦短时意图或碎片化输出。本研究提出一种基于扩散模型的框架,从与语音相关的非侵入性脑信号中解码视觉语音意图,以促进面对面的神经通信。我们设计实验,将多种音素合并训练每个音素对应的口型(viseme),旨在从神经信号中学习对应唇形表征。通过在孤立试验和连续句子上解码口型,成功重建了连贯的唇部运动,有效弥合了脑信号与动态视觉界面之间的鸿沟。结果表明,从人类神经信号中解码口型并重建说话人脸具有巨大潜力,标志着向动态神经通信系统和语音神经假体迈出关键一步。
原文摘要 · Abstract (English)
Decoding text, speech, or images from human neural signals holds promising potential both as neuroprosthesis for patients and as innovative communication tools for general users. Although neural signals contain various information on speech intentions, movements, and phonetic details, generating informative outputs from them remains challenging, with mostly focusing on decoding short intentions or producing fragmented outputs. In this study, we developed a diffusion model-based framework to decode visual speech intentions from speech-related non-invasive brain signals, to facilitate face-to-face neural communication. We designed an experiment to consolidate various phonemes to train visemes of each phoneme, aiming to learn the representation of corresponding lip formations from neural signals. By decoding visemes from both isolated trials and continuous sentences, we successfully reconstructed coherent lip movements, effectively bridging the gap between brain signals and dynamic visual interfaces. The results highlight the potential of viseme decoding and talking face reconstruction from human neural signals, marking a significant step toward dynamic neural communication systems and speech neuroprosthesis for patients.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。