用大模型模拟阅读障碍,揭示关键脑区作用。
Inducing Dyslexia in Vision Language Models
- 在视觉语言模型中找到类人字形处理单元并干扰其功能
- 干扰后模型出现读音缺陷,但整体理解能力不变
- 可复现阅读障碍者对字体敏感的特征,适合神经科学研究
阅读障碍是一种以持续阅读困难为特征的神经发育障碍,常与腹侧枕颞叶皮层的视觉词形区(VWFA)活动减弱相关。传统行为和脑成像方法虽提供重要洞见,但难以验证阅读障碍机制的因果关系。本研究利用大规模视觉语言模型(VLMs),通过识别并扰动人工模拟的词形处理单元,模拟阅读障碍。基于认知神经科学刺激,我们发现VLM中存在对视觉词形敏感的神经元单元,其响应可预测人类VWFA的神经活动。剔除这些单元后,模型在阅读任务中出现选择性受损,而一般视觉与语言理解能力保持完好。具体表现为:模型出现类似阅读障碍者的音素缺陷,但拼写处理未显著改变,并表现出对字体变化的敏感性。结果表明,该模型再现了阅读障碍的关键特征,建立了一个可用于研究脑疾病机制的计算框架。
原文摘要 · Abstract (English)
Dyslexia, a neurodevelopmental disorder characterized by persistent reading difficulties, is often linked to reduced activity of the visual word form area (VWFA) in the ventral occipito-temporal cortex. Traditional approaches to studying dyslexia, such as behavioral and neuroimaging methods, have provided valuable insights but remain limited in their ability to test causal hypotheses about the underlying mechanisms of reading impairments. In this study, we use large-scale vision-language models (VLMs) to simulate dyslexia by functionally identifying and perturbing artificial analogues of word processing. Using stimuli from cognitive neuroscience, we identify visual-word-form-selective units within VLMs and demonstrate that they predict human VWFA neural responses. Ablating model VWF units leads to selective impairments in reading tasks while general visual and language comprehension abilities remain intact. In particular, the resulting model matches dyslexic humans' phonological deficits without a significant change in orthographic processing, and mirrors dyslexic behavior in font sensitivity. Taken together, our modeling results replicate key characteristics of dyslexia and establish a computational framework for investigating brain disorders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。