用多模态联合学习+大模型,从脑电波还原高质量视频
MindCine: Multimodal EEG-to-Video Reconstruction with Large-Scale Pretrained Models
- 融合脑电与多模态信息,避免单一文本对齐偏差
- 在有限数据下实现高保真视频重建,性能超越现有方法
- 适合脑机接口、神经影像分析等研究者参考
从脑电图(EEG)信号重构人类动态视觉感知具有重要意义,因EEG具备非侵入性和高时间分辨率。然而,由于单模态依赖和数据稀缺性,现有方法仍面临挑战:1)仅将EEG与文本对齐,忽略其他模态,易过拟合;2)受限于少量EEG-视频数据,训练难收敛。为此,我们提出新框架MindCine,实现有限数据下的高保真视频重建。采用多模态联合学习策略,在训练中引入除文本外的其他模态信息,并利用预训练的大规模EEG模型缓解数据不足问题;同时设计带因果注意力的序列到序列(Seq2Seq)模型,专门解码感知信息。大量实验表明,本模型在定性和定量上均优于当前最佳方法。结果验证了不同模态互补优势的有效性,且大规模EEG模型可进一步提升性能,缓解数据限制带来的挑战。
原文摘要 · Abstract (English)
Reconstructing human dynamic visual perception from electroencephalography (EEG) signals is of great research significance since EEG's non-invasiveness and high temporal resolution. However, EEG-to-video reconstruction remains challenging due to: 1) Single Modality: existing studies solely align EEG signals with the text modality, which ignores other modalities and are prone to suffer from overfitting problems; 2) Data Scarcity: current methods often have difficulty training to converge with limited EEG-video data. To solve the above problems, we propose a novel framework MindCine to achieve high-fidelity video reconstructions on limited data. We employ a multimodal joint learning strategy to incorporate beyond-text modalities in the training stage and leverage a pre-trained large EEG model to relieve the data scarcity issue for decoding semantic information, while a Seq2Seq model with causal attention is specifically designed for decoding perceptual information. Extensive experiments demonstrate that our model outperforms state-of-the-art methods both qualitatively and quantitatively. Additionally, the results underscore the effectiveness of the complementary strengths of different modalities and demonstrate that leveraging a large-scale EEG model can further enhance reconstruction performance by alleviating the challenges associated with limited data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。