arXiv:2512.04313cs.CV2025-12被引 1

用脑电波直接生成逼真人脸表情,突破视觉依赖。

Mind-to-Face: Neural-Driven Photorealistic Avatar Synthesis via EEG Decoding

  • 通过脑电图信号映射面部几何,实现神经驱动表情生成
  • 可还原超6.5万顶点的细微表情动态,保持视角一致性
  • 适合虚拟化身、情绪交互等沉浸式应用

现有表情化系统严重依赖视觉线索,在面部遮挡或情绪内隐时失效。本文提出首个直接从非侵入性脑电图(EEG)信号解码高保真面部表情的框架——Mind-to-Face。我们构建双模态同步采集系统,在情绪刺激下获取同步的EEG与多视角面部视频,为神经到视觉学习提供精确监督。模型采用CNN-Transformer编码器将EEG信号映射为密集3D位置图,可采样超过65,000个顶点,捕捉精细几何与微妙情绪动态,并通过改进的3D高斯溅射管线进行渲染,实现逼真且视角一致的结果。大量实验表明,仅凭EEG即可可靠预测动态、个体化的面部表情,包括细微情绪反应,证明神经信号蕴含远超以往认知的情感与几何信息。Mind-to-Face建立神经驱动化身新范式,推动个性化、情绪感知的沉浸式远程交互。

原文摘要 · Abstract (English)

Current expressive avatar systems rely heavily on visual cues, failing when faces are occluded or when emotions remain internal. We present Mind-to-Face, the first framework that decodes non-invasive electroencephalogram (EEG) signals directly into high-fidelity facial expressions. We build a dual-modality recording setup to obtain synchronized EEG and multi-view facial video during emotion-eliciting stimuli, enabling precise supervision for neural-to-visual learning. Our model uses a CNN-Transformer encoder to map EEG signals into dense 3D position maps, capable of sampling over 65k vertices, capturing fine-scale geometry and subtle emotional dynamics, and renders them through a modified 3D Gaussian Splatting pipeline for photorealistic, view-consistent results. Through extensive evaluation, we show that EEG alone can reliably predict dynamic, subject-specific facial expressions, including subtle emotional responses, demonstrating that neural signals contain far richer affective and geometric information than previously assumed. Mind-to-Face establishes a new paradigm for neural-driven avatars, enabling personalized, emotion-aware telepresence and cognitive interaction in immersive environments.

脑机接口表情生成3D渲染情感计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。