用fMRI还原6秒高清流畅视频,比现有方法提升超一倍
NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction
- 分语义与感知双路径重建,兼顾画面准确性和流畅性
- 在8帧/秒下实现6秒视频还原,SSIM提升128%,时空指标提升81%
- 适合脑机接口、神经影像与视频生成研究者参考
从非侵入式脑活动fMRI中重建静态视觉刺激已取得显著进展,得益于CLIP和Stable Diffusion等深度学习模型。然而,由于连续视觉体验的时空感知解码极具挑战,当前fMRI-to-video重建研究仍较有限。我们认为,关键在于同时精准解码高层语义与底层感知流。为此,我们提出NeuroClips框架,通过语义重构器恢复视频关键帧以保证语义准确性与一致性,并利用感知重构器捕捉低层感知细节以确保视频平滑性。推理时,采用预训练的T2V扩散模型,注入关键帧与低层感知流进行视频重建。在公开fMRI-视频数据集上评估,NeuroClips可实现最高6秒、8FPS的高保真流畅视频重建,在多项指标上显著优于现有方法,如SSIM提升128%,时空指标提升81%。项目代码已开源。
原文摘要 · Abstract (English)
Reconstruction of static visual stimuli from non-invasion brain activity fMRI achieves great success, owning to advanced deep learning models such as CLIP and Stable Diffusion. However, the research on fMRI-to-video reconstruction remains limited since decoding the spatiotemporal perception of continuous visual experiences is formidably challenging. We contend that the key to addressing these challenges lies in accurately decoding both high-level semantics and low-level perception flows, as perceived by the brain in response to video stimuli. To the end, we propose NeuroClips, an innovative framework to decode high-fidelity and smooth video from fMRI. NeuroClips utilizes a semantics reconstructor to reconstruct video keyframes, guiding semantic accuracy and consistency, and employs a perception reconstructor to capture low-level perceptual details, ensuring video smoothness. During inference, it adopts a pre-trained T2V diffusion model injected with both keyframes and low-level perception flows for video reconstruction. Evaluated on a publicly available fMRI-video dataset, NeuroClips achieves smooth high-fidelity video reconstruction of up to 6s at 8FPS, gaining significant improvements over state-of-the-art models in various metrics, e.g., a 128% improvement in SSIM and an 81% improvement in spatiotemporal metrics. Our project is available at https://github.com/gongzix/NeuroClips.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。