揭示视觉语言模型中文字识别信息的传递瓶颈位置与机制。
Where Vision Becomes Text: Locating the OCR Routing Bottleneck in Vision-Language Models
- 通过因果干预定位不同架构的OCR信号关键处理层。
- 文本信号仅占总方差72.9%,且主成分可跨数据集迁移。
- 模块化模型移除OCR反而提升计数性能,说明存在干扰机制。
视觉语言模型(VLMs)能够从图像中读取文字,但文字识别(OCR)信息究竟在何处进入语言处理流程?我们针对三种架构(Qwen3-VL、Phi-4、InternVL3.5)使用因果干预方法研究了OCR路由机制。通过计算原始图像与文本掩码图像之间的激活差异,我们发现各架构存在特定的OCR瓶颈,其主导位置取决于视觉-语言融合策略:深度堆叠模型(Qwen)在中间层(约50%深度)对场景文字最敏感,而单阶段投影模型(Phi-4、InternVL)则在早期层(6–25%)达到峰值,具体层数随数据集变化。OCR信号具有极低维度特征:主成分1(PC1)可解释高达72.9%的方差。重要的是,某一数据集上学习的主成分方向可迁移至其他数据集,表明存在共享的文本处理路径。令人意外的是,在具备模块化OCR电路的模型(如Qwen3-VL-4B)中,移除OCR可使计数性能提升最高达6.9个百分点,暗示在足够模块化的架构中,OCR可能干扰其他视觉处理任务。
原文摘要 · Abstract (English)
Vision-language models (VLMs) can read text from images, but where does this optical character recognition (OCR) information enter the language processing stream? We investigate the OCR routing mechanism across three architecture families (Qwen3-VL, Phi-4, InternVL3.5) using causal interventions. By computing activation differences between original images and text-inpainted versions, we identify architecture-specific OCR bottlenecks whose dominant location depends on the vision-language integration strategy: DeepStack models (Qwen) show peak sensitivity at mid-depth (about 50%) for scene text, while single-stage projection models (Phi-4, InternVL) peak at early layers (6-25%), though the exact layer of maximum effect varies across datasets. The OCR signal is remarkably low-dimensional: PC1 captures up to 72.9% of variance. Crucially, principal component analysis (PCA) directions learned on one dataset transfer to others, demonstrating shared text-processing pathways. Surprisingly, in models with modular OCR circuits (notably Qwen3-VL-4B), OCR removal can improve counting performance (up to +6.9 percentage points), suggesting OCR interferes with other visual processing in sufficiently modular architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。