用神经引导进化生成动态视频,精准刺激大脑视觉区
NEvo: Neural-Guided Evolutionary Video Synthesis for Dynamic Visual Selectivity

- 基于动态编码模型,在结构化提示空间中进化生成视频
- 合成视频激活效果超越人工设计,揭示时间动态敏感性差异
- 适合脑科学与视觉认知研究者,可预测新实验方向
人脑通过分层组织、功能特化的区域处理动态视觉输入。尽管现有类脑编码模型能合成最优刺激以探测不同脑区的选择性,但以往研究主要局限于静态图像,动态视觉处理仍待深入。本文提出一种新型神经引导视频生成框架,可生成针对视觉皮层目标区域优化的动态刺激。方法在结构化提示空间中进行进化搜索,由动态编码模型指导,预测视频输入下的体素级响应。通过最大化目标脑区(ROI)的预测活动,框架高效发现强激活动态刺激,其效果持续优于人工设计的定位视频。合成视频恢复了腹侧、背侧和外侧通路已知的选择性,并揭示了对时间动态敏感性的系统性差异。搜索光分析进一步揭示外侧通路对复杂社会-动态特征的渐进性敏感,且在使用合成抽象非自然刺激时得到验证。本框架实现了动态视觉选择性的类脑探索,为体内实验提供新预测。
原文摘要 · Abstract (English)
The human brain processes dynamic visual input through hierarchically organized, functionally specialized regions. While recent in silico brain encoding models can synthesize optimal stimuli to probe selectivity in different brain regions, prior work has been largely limited to static images, leaving dynamic visual processing underexplored. We introduce a novel neural-guided video synthesis framework that generates stimuli optimized for target brain regions across visual cortex. Our method performs evolutionary search over a structured prompt space, guided by a dynamic encoding model that predicts voxel-level responses to video inputs. By maximizing predicted activity for a target ROI, the framework efficiently discovers hyper-activating dynamic stimuli that consistently surpass handcrafted localizer videos. The synthesized videos recover known selectivities across ventral, dorsal, and lateral pathways, and further reveal systematic differences in sensitivity to temporal dynamics. A searchlight analysis provides new insight into the progression toward increasingly complex social-dynamic features along the lateral stream, further supported by probing with synthesized abstract, non-naturalistic stimuli. Taken together, our framework enables in silico exploration of dynamic visual selectivity, with new predictions for in vivo experiments
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。