arXiv:2511.19200cs.CV2025-11

测试视觉模型能否分辨真实物体与仿制品。

Can Modern Vision Models Understand the Difference Between an Object and a Look-alike?

  • 构建了包含真实物与仿制品的RoLA数据集。
  • 发现CLIP可通过嵌入空间方向提升判别能力。
  • 适合研究模型感知与人类差异的研究者。

近期计算机视觉模型在识别任务上表现优异,但在与人类感知的对比中仍存在显著差距。一种微妙的能力是判断图像是否看起来像某物体但并非该物体实例。本文研究了如CLIP等视觉语言模型是否具备这种区分能力。我们构建了名为RoLA(Real or Lookalike)的数据集,涵盖多个类别的真实物与仿制品(如玩具、雕像、绘画、错视图像)。首先评估基于提示的基线方法,使用“真实”/“仿制”配对提示进行测试。随后估计出CLIP嵌入空间中从真实到仿制品的转换方向,并将其应用于图像和文本嵌入,提升了在Conceptual12M上的跨模态检索性能,同时改善了基于CLIP前缀的图像描述生成效果。

原文摘要 · Abstract (English)

Recent advances in computer vision have yielded models with strong performance on recognition benchmarks; however, significant gaps remain in comparison to human perception. One subtle ability is to judge whether an image looks like a given object without being an instance of that object. We study whether vision-language models such as CLIP capture this distinction. We curated a dataset named RoLA (Real or Lookalike) of real and lookalike exemplars (e.g., toys, statues, drawings, pareidolia) across multiple categories, and first evaluate a prompt-based baseline with paired "real"/"lookalike" prompts. We then estimate a direction in CLIP's embedding space that moves representations between real and lookalike. Applying this direction to image and text embeddings improves discrimination in cross-modal retrieval on Conceptual12M, and also enhances captions produced by a CLIP prefix captioner.

视觉理解多模态感知差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。