揭示视觉语言模型中多模态信息融合的局部几何路径。
Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations

- 用混合因子分析法分解激活值为局部低秩高斯邻域,捕捉细微交互结构。
- 发现不同模型在深层逐步融合或早期混合再重组的差异性融合轨迹。
- 可精准控制生成内容,性能优于传统方法,适合模型可解释性研究。
视觉语言模型(VLM)在共享残差流中处理图像块与文本标记,但两种模态间局部几何交互机制仍不清晰。现有可解释性方法多聚焦全局线性方向,可能忽略全局高维但局部低维的表示。本文提出LENS(局部邻域子空间解释),采用混合因子分析(MFA)将VLM激活分解为局部低秩高斯邻域。应用于LLaVA-1.5-7B和Qwen3-VL-8B,发现深度依赖的融合轨迹:LLaVA在后期逐层混合模态,而Qwen3-VL早期融合、部分分离后近输出层重新组合。自动化多模态标注管道为邻域赋予简洁语义标签。向邻域中心插值能因果性地引导生成,在跨模态与同模态任务中表现优异——在一项LLaVA视觉到视觉设置中,MFA得分达VL-SAE的5.7倍。人工评估显示,MFA操控效果媲美提示工程,显著优于其他基线。此外,MFA系数空间使Qwen3-VL在最深层图像到渲染文本检索任务中,R@1从14.9%提升至48.6%。消融实验表明融合轨迹对组件数、局部秩及模态纯度阈值均稳定。结果支持局部几何邻域作为可解释且因果的交叉模态表示单位。
原文摘要 · Abstract (English)
Vision-language models (VLMs) process image patches and text tokens in a shared residual stream, but the local geometry through which the two modalities interact remains poorly understood. Most interpretability methods identify global linear directions, which may miss representations that are globally high-dimensional but locally low-dimensional. We introduce LENS (Local Explanation of Neighborhood Subspaces), a method that decomposes VLM activations into local low-rank Gaussian neighborhoods using a Mixture of Factor Analyzers. Applied to LLaVA-1.5-7B and Qwen3-VL-8B, LENS reveals distinct depth-dependent fusion trajectories consistent with each model's fusion mechanism: LLaVA progressively mixes modalities at later layers, whereas Qwen3-VL mixes them early, partially re-segregates them, and recombines them near the output. An automated multimodal labeling pipeline assigns concise semantic descriptions to these neighborhoods. Interpolating activations toward neighborhood centroids causally redirects generation within and across modalities and outperforms difference-in-means and VL-SAE in most evaluated conditions; in one LLaVA vision-to-vision setting, MFA achieves 5.7 times the VL-SAE score. Human evaluation finds MFA steering competitive with prompting and substantially stronger than the other intervention baselines. Finally, the MFA coefficient space improves Qwen3-VL image-to-rendered-text retrieval at the deepest evaluated layer from 14.9% to 48.6% R@1. Ablations show that the reported fusion trajectories are stable across component counts, local ranks, and modality-purity thresholds. These results support local geometric neighborhoods as useful interpretable and causal units for analyzing cross-modal representations in the evaluated VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。