arXiv:2607.22302cs.CV2026-07

用fMRI重建高清人脸视频,首次实现动态表情与身份的精准还原。

fMRI2Face: A Full-HD fMRI-Video Dataset and Geometry-Guided Neural Decoding Framework for Dynamic Human Face Reconstruction

论文配图:fMRI2Face: A Full-HD fMRI-Video Dataset and Geometry-Guided Neural Decoding Framework for Dynamic Human Face Reconstruction
图 1 · 摘自论文原文
  • 基于脑活动生成人脸视频,融合几何引导与外观控制
  • 在1920×1080分辨率下实现62,856对齐样本重建
  • 适合脑机接口、面部感知研究者参考

从脑活动重建动态人脸为研究认知中身份、表情与面部运动提供了强大手段。然而,现有进展受限于稀缺的可控高分辨率神经数据集及难以同时恢复身份特征与动态变化的方法。我们提出fMRI-Face,首个与可控全高清(1920×1080)数字人脸视频配对的fMRI数据集。参与者在扫描中观看无背景、受控身份、表情和头部姿态的逼真人脸视频,同步采集fMRI信号。数据集包含62,856个成对样本,为动态人脸感知与重建提供结构化资源。基于此,我们提出fMRI2Face——一种几何引导的神经视频解码框架,从fMRI信号重建人脸视频。该框架通过脑活动提取两个互补控制:脑源外观上下文(捕获身份相关视觉属性)与可变形3D面部控制(提供姿态、表情与非刚性动态的显式几何引导)。二者通过辅助潜在空间补全的神经控制视频扩散模型融合,实现高保真人脸视频重建。实验表明,fMRI2Face在重建保真度、身份保留、面部几何与运动一致性方面均显著优于主流基线。fMRI-Face与fMRI2Face共同建立了一个可控的研究平台,为动态人脸感知与基于fMRI的数字人重建提供新基准。

原文摘要 · Abstract (English)

Reconstructing dynamic human faces from brain activity provides a powerful way to study how the mind perceives identity, expression, and facial motion. However, progress in fMRI-based face decoding has been limited by scarce controlled, high-resolution neural datasets and by methods that struggle to recover both identity-specific appearance and time-varying facial dynamics. We present fMRI-Face, the first fMRI dataset paired with controllable full-HD digital human facial videos rendered at 1920$\times$1080 resolution. During scanning, participants watched photorealistic, background-free facial videos with controlled identity, expression, and head pose, while fMRI activity was recorded. The resulting dataset contains 62,856 paired fMRI-video samples, providing a structured resource for studying dynamic face perception and reconstruction. Building on this dataset, we propose fMRI2Face, a geometry-guided neural video decoding framework for reconstructing facial videos from fMRI signals. fMRI2Face derives two complementary neural controls from brain activity: Brain-derived Appearance Context, which captures global identity-related visual attributes, and Morphable 3D Facial Control, which provides explicit geometry-aware guidance for pose, expression, and non-rigid facial dynamics. These controls are integrated through Neural-Controlled Video Diffusion with auxiliary latent completion, enabling high-fidelity facial video reconstruction directly from brain activity. Experiments show that fMRI2Face consistently improves reconstruction fidelity, identity preservation, facial geometry, and motion consistency over representative neural decoding baselines. Together, fMRI-Face and fMRI2Face establish a controlled platform for studying dynamic face perception and provide a new benchmark for fMRI-based digital human reconstruction.

脑机接口人脸重建fMRI视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。