用大模型解码脑电数据,实现多片段视频重建。
MindShot: Multi-Shot Video Reconstruction from fMRI with LLM Decoding
- 分段预测+大模型生成关键帧描述,突破单片段限制。
- 重建精度超越现有方法,语义相似度提升71.8%。
- 适合脑机接口与认知神经科学研究者。
从fMRI中重建动态视频对理解视觉认知和实现生动的脑机接口至关重要。然而,现有方法仅限于单片段重建,无法应对真实场景中的多片段特性。多片段重建面临三大挑战:不同片段的fMRI信号混合、fMRI与视频的时间分辨率不匹配导致快速场景变化模糊,以及缺乏专门的多片段fMRI-视频数据集。为此,我们提出一种新型的分而解之框架,用于多片段fMRI视频重建。核心创新包括:(1) 射门边界预测模块,显式将混合的fMRI信号分解为片段特定段;(2) 基于大模型的生成式关键帧描述,从每段信号中解码鲁棒的文本描述,通过高层语义克服时间模糊;(3) 从现有数据集合成大规模新数据(20,000样本)。实验表明,本框架在多片段重建保真度上优于最先进方法。消融实验确认信号分解与语义提示的关键作用,其中分解使解码描述的CLIP相似度提升71.8%。该工作建立了多片段fMRI重建的新范式,通过显式分解与语义提示,实现复杂视觉叙事的准确恢复。
原文摘要 · Abstract (English)
Reconstructing dynamic videos from fMRI is important for understanding visual cognition and enabling vivid brain-computer interfaces. However, current methods are critically limited to single-shot clips, failing to address the multi-shot nature of real-world experiences. Multi-shot reconstruction faces fundamental challenges: fMRI signal mixing across shots, the temporal resolution mismatch between fMRI and video obscuring rapid scene changes, and the lack of dedicated multi-shot fMRI-video datasets. To overcome these limitations, we propose a novel divide-and-decode framework for multi-shot fMRI video reconstruction. Our core innovations are: (1) A shot boundary predictor module explicitly decomposing mixed fMRI signals into shot-specific segments. (2) Generative keyframe captioning using LLMs, which decodes robust textual descriptions from each segment, overcoming temporal blur by leveraging high-level semantics. (3) Novel large-scale data synthesis (20k samples) from existing datasets. Experimental results demonstrate our framework outperforms state-of-the-art methods in multi-shot reconstruction fidelity. Ablation studies confirm the critical role of fMRI decomposition and semantic captioning, with decomposition significantly improving decoded caption CLIP similarity by 71.8%. This work establishes a new paradigm for multi-shot fMRI reconstruction, enabling accurate recovery of complex visual narratives through explicit decomposition and semantic prompting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。