arXiv:2509.18177cs.CVcs.AI2025-09

用新框架生成数据集,检验AI对位置关系的理解能力

A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts

论文配图:A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
图 1 · 摘自论文原文
  • 构建Scrapbook框架,生成带语言多样性的概念探测数据
  • 现有模型对位置信息理解差,MobileVLM-V2答错率高
  • 适合研究AI基础认知能力与模型鲁棒性的人看

本文提出Scrapbook框架,一种用于生成大规模数据集的新方法,旨在探测人工智能模型对基本概念的习得情况。该框架聚焦于物体识别、绝对与相对位置、属性识别等核心概念。通过生成大量关于单一概念的问题及广泛的语言变体,旨在在处理复杂任务前验证模型对这些基础元素的理解。实验结果表明,尽管当前模型在物体识别和计数方面表现良好,但在理解位置信息和应对附加约束问题时存在困难。MobileVLM-V2模型出现显著答案不一致和合理错误回答,其他模型则表现出肯定倾向,且在几何形状与位置相关问题上表现不佳,反映出在理解一致性方面仍有改进空间。该框架为系统评估和提升AI模型性能提供了有力工具。

原文摘要 · Abstract (English)

In this paper, we present the Scrapbook framework, a novel methodology designed to generate extensive datasets for probing the learned concepts of artificial intelligence (AI) models. The framework focuses on fundamental concepts such as object recognition, absolute and relative positions, and attribute identification. By generating datasets with a large number of questions about individual concepts and a wide linguistic variation, the Scrapbook framework aims to validate the model's understanding of these basic elements before tackling more complex tasks. Our experimental findings reveal that, while contemporary models demonstrate proficiency in recognizing and enumerating objects, they encounter challenges in comprehending positional information and addressing inquiries with additional constraints. Specifically, the MobileVLM-V2 model showed significant answer disagreements and plausible wrong answers, while other models exhibited a bias toward affirmative answers and struggled with questions involving geometric shapes and positional information, indicating areas for improvement in understanding and consistency. The proposed framework offers a valuable instrument for generating diverse and comprehensive datasets, which can be utilized to systematically assess and enhance the performance of AI models.

概念探测位置理解数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。