发现多模态模型在空间推理中被词汇误导,通过轻量微调可显著提升准确性。
Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning

- 识别出模型因空间词汇干扰而误选答案的机制性偏差
- 仅用少量合成数据微调,使四分类准确率最高提升100点
- 适用于改进视觉-语言模型的空间推理能力,尤其关注错误归因
多模态大语言模型在空间选择题上仍不可靠,其失败常归因于视觉信息关注不足。我们发现一种互补性失效模式:空间关系词作为语义干扰项,会诱使模型偏向对应选项。在九个开源权重的MLLM上验证了该现象的普遍性。进一步识别出‘二元稳定但三元脆弱’案例——模型能正确回答二选一问题,却在新增第三个选项时总选错。通过机制可解释性工具分析,发现故障源于语言侧而非视觉侧:视觉注意力与残差流探测显示正确空间关系仍被保留;而无关选项控制、激活修补及稀疏组件干预表明偏差源自特定的语言模型通道与神经元。据此,仅用极小量单对象对合成数据进行轻量级DPO微调,即可缓解偏差,在合成数据上使四分类鲁棒准确率最高提升100点,在WhatsUp、SpatialMQA-Direct和VSR等更广评估集上分别提升68.0、32.6和20.1点。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) remain unreliable on spatial multiple-choice questions, and their failures are often attributed to poorly attended visual information. We identify a complementary failure mode, spatial lexical bias: a spatial relation word added to the answer options can act as a lexical-semantic distractor that draws the model's decision toward that option. Using nine open-weight MLLMs, we show that this phenomenon is widespread. We then isolate diagnostic cases in which a model answers a binary spatial question correctly yet consistently chooses a newly added third spatial option, which we call binary-stable but ternary-fragile cases. Leveraging mechanistic interpretability tools on these cases, we find that the failure arises on the language side rather than the visual side: visual attention analyses and residual-stream probes show the correct spatial relation remains internally available, while irrelevant-option controls, activation patching, and sparse component interventions trace the bias to specific LLM-side channels and neurons. Accordingly, we show that a lightweight LLM-only DPO update on tiny single-object-pair synthetic data mitigates the bias, lifting four-way robust accuracy by up to 100 points on synthetic data, and by 68.0, 32.6, and 20.1 points on broader evaluation datasets WhatsUp, SpatialMQA-Direct, and VSR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。