发现视觉语言模型具备人类冲突适应行为,且机制与人脑相似。
Conflict Adaptation in Vision-Language Models
- 通过序列斯特鲁普任务测试,12/13模型表现出冲突适应现象。
- 在InternVL 3.5 4B中识别出文本与颜色共用的超节点,其大小反映自动性差异。
- 定位到第24-25层的冲突调节超节点,移除后显著增加错误率。
人类认知控制的一个特征是冲突适应:在高冲突试次后,后续高冲突试次的表现会提升。该现象解释了稀缺的认知控制资源如何被调动。我们使用序列斯特鲁普任务发现,在测试的13个视觉语言模型(VLMs)中,12个表现出与冲突适应一致的行为,唯一例外可能因性能接近天花板。为理解该行为的表征基础,我们采用稀疏自编码器(SAEs)分析InternVL 3.5 4B模型,发现文本与颜色在早期和晚期层均存在部分重叠的超节点,其相对大小对应人类阅读与颜色命名的自动性差异。此外,我们在第24-25层识别出一个受冲突调节的超节点,其移除显著增加斯特鲁普错误率,而对一致试次影响甚微。
原文摘要 · Abstract (English)
A signature of human cognitive control is conflict adaptation: improved performance on a high-conflict trial following another high-conflict trial. This phenomenon offers an account for how cognitive control, a scarce resource, is recruited. Using a sequential Stroop task, we find that 12 of 13 vision-language models (VLMs) tested exhibit behavior consistent with conflict adaptation, with the lone exception likely reflecting a ceiling effect. To understand the representational basis of this behavior, we use sparse autoencoders (SAEs) to identify task-relevant supernodes in InternVL 3.5 4B. Partially overlapping supernodes emerge for text and color in both early and late layers, and their relative sizes mirror the automaticity asymmetry between reading and color naming in humans. We further isolate a conflict-modulated supernode in layers 24-25 whose ablation significantly increases Stroop errors while minimally affecting congruent trials.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。