arXiv:2501.15144cs.CV2025-01中稿 · DICTA 2025被引 2

研究视觉语言模型对基础图形的感知能力,发现输出格式影响模型表现。

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models

  • 用控制变量的2D图形测试模型理解能力,评估空间位置、遮挡等属性
  • 句子输出比元组格式在跨领域任务中表现更优,尤其在差异大时
  • 数值标记加权损失能提升空间测量精度,适合需要精准推理的场景

本研究针对当前视觉语言模型(VLMs)在基础图形视觉理解与属性测量方面的能力,构建了一个聚焦于受控2D形状配置的基准测试,涵盖空间位置、遮挡、旋转、大小、形状类型、象限、中心坐标、旋转角度、遮挡状态及颜色等属性的变化。我们使用低秩适应(LoRA)微调了参数量为2B-8B的先进VLMs,并在所提基准中的多个域外(OD)场景下验证其性能。结果表明,连贯的句子输出优于元组格式,尤其在领域差异较大的情形下;同时,通过在损失计算中对数值标记进行缩放,可显著增强模型的数值近似能力,进一步提升空间与测量任务的表现。这些发现强调了输出格式设计、损失缩放策略以及鲁棒泛化技术在提升VLM训练与微调效果中的重要性,尤其是在需要精确空间近似和强域外泛化能力的任务中。

原文摘要 · Abstract (English)

This work investigates the capabilities of current vision-language models (VLMs) in visual understanding and attribute measurement of primitive shapes using a benchmark focused on controlled 2D shape configurations with variations in spatial positioning, occlusion, rotation, size, and shape attributes such as type, quadrant, center-coordinates, rotation, occlusion status, and color as shown in Figure 1 and supplementary Figures S3-S81. We fine-tune state-of-the-art VLMs (2B-8B parameters) using Low-Rank Adaptation (LoRA) and validate them on multiple out-of-domain (OD) scenarios from our proposed benchmark. Our findings reveal that coherent sentence-based outputs outperform tuple formats, particularly in OD scenarios with large domain gaps. Additionally, we demonstrate that scaling numeric tokens during loss computation enhances numerical approximation capabilities, further improving performance on spatial and measurement tasks. These results highlight the importance of output format design, loss scaling strategies, and robust generalization techniques in enhancing the training and fine-tuning of VLMs, particularly for tasks requiring precise spatial approximations and strong OD generalization.

视觉语言模型空间理解输出格式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。