小模型视觉能力下降比推理弱更严重,提出提取+思考新方法提升性能。
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- 通过分离分析发现:缩小语言模型会严重损伤视觉感知能力。
- 在多个数据集上,小模型的视觉表现下降幅度超过推理能力退化。
- 提出提取+思考框架,显著提升小模型在视觉任务上的效率与准确率。
尽管大规模多模态模型在视觉理解与推理方面取得显著进展,但实际应用需求推动对小型高效系统的探索。本文系统分析了多模态模型在缩小规模时智能表现的变化,重点考察小语言模型(LLM)容量降低对多模态能力的影响。初步结果表明,模型缩放对视觉能力的损害远超其对语言模型原有能力的影响。进一步分析发现,这种性能下降并非仅源于推理能力减弱,而是深层的感知能力损失。为此,我们提出视觉提取调优(visual extraction tuning),通过显式训练模型稳定提取指令相关的视觉细节,并结合逐步推理生成答案。该方法构成「Extract+Think」框架,在保证高效的同时显著提升小模型性能,为小型多模态系统提供新范式。
原文摘要 · Abstract (English)
Scaling up multimodal models has enabled remarkable advances in visual understanding and reasoning, but practical demands call for smaller, efficient systems. In this work, we conduct a principled analysis of downscaling intelligence in multimodal models, examining how reduced large language model (LLM) capacity affects multimodal capabilities. Our initial findings reveal an interesting trend: LLM downscaling disproportionately affects visual capabilities, rather than abilities inherited from the LLM. We then examine whether this drop mainly reflects the expected decline in visual reasoning or a more fundamental loss of perceptual abilities. Isolating the effect of LLM downscaling on perception, we find performance still drops sharply, often matching or exceeding the impact on reasoning. To address this bottleneck, we introduce visual extraction tuning, which explicitly trains the model to extract instruction-relevant visual details consistently across tasks. With these extracted visual details, we then apply step-by-step reasoning to generate answers. Together, these components form our Extract+Think approach, setting a new standard for efficiency and performance in this space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。