arXiv:2412.18224cs.CVcs.AI2024-12被引 4

用扩散模型扩展视觉位置数据,提升大模型对空间关系的感知能力

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules

  • 通过扩散模型可控扩充图像位置数据,融合多种视觉编码器增强感知
  • 模型在空间推理测试集上准确率提升超27%,兼顾指令泛化与位置区分
  • 适合研究视觉语言模型、空间认知或需要精准位置理解的场景

区分空间关系是人类认知的基础,需对跨实例进行细粒度感知。尽管现有基准如MME、MMBench和SEED已涵盖视觉空间推理(VSR)能力评估,但针对视觉位置推理的高质量、大规模训练与优化数据仍不足。为此,我们首先诊断了当前视觉大语言模型(VLLMs)在VSR数据集上的表现,提出统一测试集。发现其存在对语言指令过度敏感、对视觉位置信息敏感度不足的矛盾。通过从调优数据和模型结构两方面扩展基准,缓解该问题。首次使用扩散模型可控地扩展空间定位图像数据,并将原始视觉编码器(CLIP)与SigLIP、SAM、DINO三种强大编码器结合。在数据与模型规模组合实验后,构建出具备空间推理专长的VLLM(VSRE),不仅更适应不同指令,还能准确区分视觉位置差异。VSRE在VSR测试集上准确率提升超过27%,在该数据集及其它基准相关子集上均表现优异。项目已开源数据与模型,链接:https://github.com/peijin360/vsre,旨在推动VLLM在视觉空间推理学习中的进展。

原文摘要 · Abstract (English)

Distinguishing spatial relations is a basic part of human cognition which requires fine-grained perception on cross-instance. Although benchmarks like MME, MMBench and SEED comprehensively have evaluated various capabilities which already include visual spatial reasoning(VSR). There is still a lack of sufficient quantity and quality evaluation and optimization datasets for Vision Large Language Models(VLLMs) specifically targeting visual positional reasoning. To handle this, we first diagnosed current VLLMs with the VSR dataset and proposed a unified test set. We found current VLLMs to exhibit a contradiction of over-sensitivity to language instructions and under-sensitivity to visual positional information. By expanding the original benchmark from two aspects of tunning data and model structure, we mitigated this phenomenon. To our knowledge, we expanded spatially positioned image data controllably using diffusion models for the first time and integrated original visual encoding(CLIP) with other 3 powerful visual encoders(SigLIP, SAM and DINO). After conducting combination experiments on scaling data and models, we obtained a VLLM VSR Expert(VSRE) that not only generalizes better to different instructions but also accurately distinguishes differences in visual positional information. VSRE achieved over a 27\% increase in accuracy on the VSR test set. It becomes a performant VLLM on the position reasoning of both the VSR dataset and relevant subsets of other evaluation benchmarks. We open-sourced the expanded model with data and Appendix at \url{https://github.com/peijin360/vsre} and hope it will accelerate advancements in VLLM on VSR learning.

视觉推理空间关系大模型扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。