arXiv:2512.18951cs.LG2025-12

对比婴儿与网络训练模型对物体属性的识别能力,发现各有优劣。

Benchmarking Attribute Discrimination in Infant-Scale Vision-Language Models

  • 用合成图像控制颜色、大小、纹理,构建可区分属性的测试集
  • 婴儿模型在大小和纹理上表现好,但颜色视觉识别差,文本引导下更弱
  • 网络模型能很好理解颜色文字,但对大小的视觉辨别力较弱

婴儿从有限经验中学习物体类别及颜色、大小、纹理等细粒度视觉属性。此前的婴儿规模视觉-语言模型主要评估对象识别能力,未考察类内属性区分能力。本文提出一个受控基准,通过合成渲染在67种日常物体类别中独立调节颜色、大小和纹理,以分离属性值与物体身份。在图像仅原型测试和带属性-物体提示的图文测试两种设置下,评估婴儿训练模型(CVCL和婴儿训练DINO基线)与网络规模模型(CLIP、SigLIP、ResNeXt)。结果发现视觉与语言属性信息存在解耦:婴儿模型对大小形成强视觉表征,纹理辨别能力与其它模型相当,但视觉颜色识别表现差;在图文任务中难以锚定颜色,仅表现出适度的大小锚定。相比之下,网络训练模型能从文本中强地锚定颜色,但在视觉大小辨别上表现较弱。

原文摘要 · Abstract (English)

Infants learn not only object categories but also fine-grained visual attributes such as color, size, and texture from limited experience. Prior infant-scale vision--language models have mainly been evaluated on object recognition, leaving open whether they support within-class attribute discrimination. We introduce a controlled benchmark that varies color, size, and texture across 67 everyday object classes using synthetic rendering to decouple attribute values from object identity. We evaluate infant-trained models (CVCL and an infant-trained DINO baseline) against web-scale and ImageNet models (CLIP, SigLIP, ResNeXt) under two complementary settings: an image-only prototype test and a text--vision test with attribute--object prompts. We find a dissociation between visual and linguistic attribute information: infant-trained models form strong visual representations for size and discriminate texture comparably to other models, but perform poorly on visual color discrimination, and in the text--vision setting they struggle to ground color and show only modest size grounding. In contrast, web-trained vision--language models strongly ground color from text while exhibiting weaker visual size discrimination.

视觉-语言属性识别婴儿模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。