arXiv:2602.00108cs.CVcs.AI2026-02中稿 · ICCV被引 2

构建可控空间关系的合成计数数据集,提升视觉语言模型泛化能力

SITUATE -- Synthetic Object Counting Dataset for VLM training

  • 基于可控遮挡与空间布局生成合成图像,精准控制计数场景
  • 在Pixmo测试集上微调后准确率提升,证明对分布外数据泛化有效
  • 适合训练需空间理解的视觉语言模型,尤其关注计数任务的研究者

我们提出SITUATE,一个专为视觉语言模型计数任务设计的新颖数据集,具有空间约束。该数据集填补了简单2D数据集(如VLMCountBench)与现实世界中模糊且不可控的数据集(如TallyQA)之间的空白,能精确控制遮挡与空间组合。实验表明,将Qwen VL 2.5 7B在SITUATE上微调后,可提升其在Pixmo count测试集上的准确率,而反向则无效。通过与其他主流计数基准对比,并与等规模的Pixmo count子集进行微调对比,验证了SITUATE的有效性。

原文摘要 · Abstract (English)

We present SITUATE, a novel dataset designed for training and evaluating Vision Language Models on counting tasks with spatial constraints. The dataset bridges the gap between simple 2D datasets like VLMCountBench and often ambiguous real-life datasets like TallyQA, which lack control over occlusions and spatial composition. Experiments show that our dataset helps to improve generalization for out-of-distribution images, since a finetune of Qwen VL 2.5 7B on SITUATE improves accuracy on the Pixmo count test data, but not vice versa. We cross validate this by comparing the model performance across established other counting benchmarks and against an equally sized fine-tuning set derived from Pixmo count.

视觉语言模型合成数据计数任务空间约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。