用脑电波重建动态视频,提升清晰度与时间连贯性。
DynaMind: Reconstructing Dynamic Visual Scenes from EEG by Aligning Temporal Dynamics and Multimodal Semantics to Guided Diffusion
- 通过区域感知与时间对齐模块,融合脑区语义与神经动态
- 在SEED-DV数据集上视频准确率提升12.5%,像素级质量提升9.4% SSIM
- 适合脑机接口、视觉重建与神经编码研究者参考
从脑电图(EEG)信号中重构动态视觉场景仍是脑解码的核心挑战,受限于EEG低空间分辨率、神经记录与视频动态间的时间错位,以及脑活动中语义信息利用不足。现有方法常难以同时保证动态连贯性与复杂语义上下文。为此,我们提出DynaMind框架,通过三个核心模块联合建模神经动态与语义特征:区域感知语义映射器(RSM)、时间感知动态对齐器(TDA)和双引导视频重构器(DGVR)。RSM利用区域感知编码器提取不同脑区的多模态语义特征,并聚合为统一扩散先验;TDA生成动态潜在序列(即蓝图),确保特征表示与原始神经记录的时间一致性。在此基础上,DGVR以语义扩散先验为指导,将时间感知蓝图转化为高保真视频重建。在SEED-DV数据集上,DynaMind达到新SOTA,视频与帧级准确率分别提升12.5和10.3个百分点;像素级质量显著改善,SSIM提升9.4%,FVMD降低19.7%。该工作实现了神经动态与高保真视觉语义之间的关键衔接。
原文摘要 · Abstract (English)
Reconstruction dynamic visual scenes from electroencephalography (EEG) signals remains a primary challenge in brain decoding, limited by the low spatial resolution of EEG, a temporal mismatch between neural recordings and video dynamics, and the insufficient use of semantic information within brain activity. Therefore, existing methods often inadequately resolve both the dynamic coherence and the complex semantic context of the perceived visual stimuli. To overcome these limitations, we introduce DynaMind, a novel framework that reconstructs video by jointly modeling neural dynamics and semantic features via three core modules: a Regional-aware Semantic Mapper (RSM), a Temporal-aware Dynamic Aligner (TDA), and a Dual-Guidance Video Reconstructor (DGVR). The RSM first utilizes a regional-aware encoder to extract multimodal semantic features from EEG signals across distinct brain regions, aggregating them into a unified diffusion prior. In the mean time, the TDA generates a dynamic latent sequence, or blueprint, to enforce temporal consistency between the feature representations and the original neural recordings. Together, guided by the semantic diffusion prior, the DGVR translates the temporal-aware blueprint into a high-fidelity video reconstruction. On the SEED-DV dataset, DynaMind sets a new state-of-the-art (SOTA), boosting reconstructed video accuracies (video- and frame-based) by 12.5 and 10.3 percentage points, respectively. It also achieves a leap in pixel-level quality, showing exceptional visual fidelity and temporal coherence with a 9.4% SSIM improvement and a 19.7% FVMD reduction. This marks a critical advancement, bridging the gap between neural dynamics and high-fidelity visual semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。