arXiv:2603.18523cs.CVcs.AI2026-03被引 2

揭示视觉语言模型的计数机制并提升其推理能力

Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models

  • 通过合成数据与机械分析,发现模型有类人计数行为
  • 仅微调计数任务,即实现+8.36%跨分布性能提升
  • 适合想提升模型视觉推理能力的研究者使用

计数是检验大视觉语言模型(LVLM)推理能力的简单而有力指标,要求模型识别每个独立物体并累加。本研究通过受控的合成与真实世界基准,结合机械可解释性分析,探究了LVLM如何实现计数。结果表明,LVLM在小数量下表现精确,在大数量时呈现噪声估计,具有类人计数行为。我们提出两种新可解释性方法:视觉激活修补和HeadLens,揭示了跨多种视觉推理任务共享的“计数电路”。基于此,提出一种轻量级干预策略,仅用大量可用的合成图像对任意预训练LVLM进行计数微调。尽管微调范围狭窄,该方法不仅提升了分布内合成数据的计数准确率,还在分布外计数基准上平均提高+8.36%,并在复杂通用视觉推理任务中使Qwen2.5-VL平均提升+1.54%。这些发现凸显计数在视觉推理中的核心作用,提示通过针对性增强计数机制可有效提升整体视觉推理能力。

原文摘要 · Abstract (English)

Counting serves as a simple but powerful test of a Large Vision-Language Model's (LVLM's) reasoning; it forces the model to identify each individual object and then add them all up. In this study, we investigate how LVLMs implement counting using controlled synthetic and real-world benchmarks, combined with mechanistic analyses. Our results show that LVLMs display a human-like counting behavior, with precise performance on small numerosities and noisy estimation for larger quantities. We introduce two novel interpretability methods, Visual Activation Patching and HeadLens, and use them to uncover a structured "counting circuit" that is largely shared across a variety of visual reasoning tasks. Building on these insights, we propose a lightweight intervention strategy that exploits simple and abundantly available synthetic images to fine-tune arbitrary pretrained LVLMs exclusively on counting. Despite the narrow scope of this fine-tuning, the intervention not only enhances counting accuracy on in-distribution synthetic data, but also yields an average improvement of +8.36% on out-of-distribution counting benchmarks and an average gain of +1.54% on complex, general visual reasoning tasks for Qwen2.5-VL. These findings highlight the central, influential role of counting in visual reasoning and suggest a potential pathway for improving overall visual reasoning capabilities through targeted enhancement of counting mechanisms.

视觉推理可解释性计数机制微调策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。