用深度模型预测大脑对增强图像的反应,验证其泛化能力。
Generalizability analysis of deep learning predictions of human brain responses to augmented and semantically novel visual stimuli
- 构建脑编码模型,评估图像增强对视觉皮层的影响
- 模型能准确预测特定物体增强后的神经响应
- 适用于AR/VR和模型驱动的设计优化
本研究旨在评估基于神经网络的方法在探索图像增强技术对视觉皮层激活影响方面的有效性与可靠性。我们选取了参与2023年Algonauts项目挑战赛前10名的先进脑编码模型,分析其对不同图像增强技术引发神经反应的预测能力。由于脑成像数据获取成本高昂,无法获得真实数据,因此通过一系列实验进行验证。具体而言,我们考察模型对已知影响特定脑区的物体(如人脸和文字)增强后的反应预测能力,并研究模型对训练中未见的语义外分布刺激的激活预测表现。结果表明,所提框架中的模型具有良好的泛化能力,有望用于识别特定任务下的最优视觉增强滤波器、实现模型驱动的设计策略,以及支持AR/VR应用。
原文摘要 · Abstract (English)
The purpose of this work is to investigate the soundness and utility of a neural network-based approach as a framework for exploring the impact of image enhancement techniques on visual cortex activation. In a preliminary study, we prepare a set of state-of-the-art brain encoding models, selected among the top 10 methods that participated in The Algonauts Project 2023 Challenge [16]. We analyze their ability to make valid predictions about the effects of various image enhancement techniques on neural responses. Given the impossibility of acquiring the actual data due to the high costs associated with brain imaging procedures, our investigation builds up on a series of experiments. Specifically, we analyze the ability of brain encoders to estimate the cerebral reaction to various augmentations by evaluating the response to augmentations targeting objects (i.e., faces and words) with known impact on specific areas. Moreover, we study the predicted activation in response to objects unseen during training, exploring the impact of semantically out-of-distribution stimuli. We provide relevant evidence for the generalization ability of the models forming the proposed framework, which appears to be promising for the identification of the optimal visual augmentation filter for a given task, model-driven design strategies as well as for AR and VR applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。