用AI生成5400张真实感图像,挑战模型对细微视觉概念的推理能力
Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems
- 用视觉语言模型流水线生成真实感图像,构建大规模细粒度BP数据集
- 5400个实例验证,主流VLM在细粒度概念识别上表现不佳
- 适合研究视觉推理局限性、模型泛化能力的学者参考
Bongard问题(BPs)为抽象视觉推理(AVR)提供了一个挑战性测试平台,要求模型从少数示例中识别视觉概念并用自然语言描述。早期的BP基准使用黑白合成图,难以反映真实场景复杂性;后续数据集虽采用真实图像,但概念仅依赖高层特征即可识别,降低了任务难度。相比之下,最近发布的Bongard-RWR数据集尝试用细粒度真实图像再现原问题中的抽象概念,但其人工构造导致仅有60个实例,限制了评估稳健性。本文提出Bongard-RWR+,一个包含5,400个实例的全新数据集,通过视觉语言模型(VLM)管道生成类真实世界图像来表示原始BP中的抽象概念。我们利用Pixtral-12B为人工筛选图像生成描述,并基于这些描述用Flux.1-dev合成新图像,经人工验证确保生成图像忠实反映目标概念。我们在多种BP形式下评估最先进的VLM,包括二分类与多分类以及文本答案生成。结果表明,尽管VLM能识别粗粒度视觉概念,但在细粒度概念辨识上持续表现不佳,揭示其推理能力的局限性。
原文摘要 · Abstract (English)
Bongard Problems (BPs) provide a challenging testbed for abstract visual reasoning (AVR), requiring models to identify visual concepts fromjust a few examples and describe them in natural language. Early BP benchmarks featured synthetic black-and-white drawings, which might not fully capture the complexity of real-world scenes. Subsequent BP datasets employed real-world images, albeit the represented concepts are identifiable from high-level image features, reducing the task complexity. Differently, the recently released Bongard-RWR dataset aimed at representing abstract concepts formulated in the original BPs using fine-grained real-world images. Its manual construction, however, limited the dataset size to just $60$ instances, constraining evaluation robustness. In this work, we introduce Bongard-RWR+, a BP dataset composed of $5\,400$ instances that represent original BP abstract concepts using real-world-like images generated via a vision language model (VLM) pipeline. Building on Bongard-RWR, we employ Pixtral-12B to describe manually curated images and generate new descriptions aligned with the underlying concepts, use Flux.1-dev to synthesize images from these descriptions, and manually verify that the generated images faithfully reflect the intended concepts. We evaluate state-of-the-art VLMs across diverse BP formulations, including binary and multiclass classification, as well as textual answer generation. Our findings reveal that while VLMs can recognize coarse-grained visual concepts, they consistently struggle with discerning fine-grained concepts, highlighting limitations in their reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。