arXiv:2411.09211cs.AI2024-11被引 2

用脑电信号实时还原说话时的口型动态,实现自然交流

Dynamic Neural Communication: Convergence of Computer Vision and Brain-Computer Interface

论文配图:Dynamic Neural Communication: Convergence of Computer Vision and Brain-Computer Interface
图 1 · 摘自论文原文
  • 通过神经信号捕捉说话意图,分步解码口形变化
  • 能快速重建自然发音时的唇部运动动态序列
  • 适合脑机接口与视觉交互研究者参考

将人类神经信号解析为静态语言意图(如文字或图像)和动态语言意图(如音频或视频)正展现出作为创新沟通工具的巨大潜力。人类交流伴随多种特征,如发音动作、面部表情和内在言语,这些均反映在神经信号中。然而,现有研究多生成短片段输出,如何融合神经信号中的多种特征实现信息丰富且连贯的交流仍具挑战。本研究提出一种动态神经通信方法,结合计算机视觉与脑机接口技术,从神经信号中捕捉用户意图,并在短时间步内解码口形(visemes),生成动态视觉输出。结果表明,该方法可快速捕获并重构自然发音尝试时的唇部运动,实现了计算机视觉与脑机接口融合下的动态神经通信。

原文摘要 · Abstract (English)

Interpreting human neural signals to decode static speech intentions such as text or images and dynamic speech intentions such as audio or video is showing great potential as an innovative communication tool. Human communication accompanies various features, such as articulatory movements, facial expressions, and internal speech, all of which are reflected in neural signals. However, most studies only generate short or fragmented outputs, while providing informative communication by leveraging various features from neural signals remains challenging. In this study, we introduce a dynamic neural communication method that leverages current computer vision and brain-computer interface technologies. Our approach captures the user's intentions from neural signals and decodes visemes in short time steps to produce dynamic visual outputs. The results demonstrate the potential to rapidly capture and reconstruct lip movements during natural speech attempts from human neural signals, enabling dynamic neural communication through the convergence of computer vision and brain--computer interface.

脑机接口动态生成神经解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。