arXiv:2607.02402cs.CV2026-07

让AI从图片集推断共同概念并生成符合该概念的新图

Show Me Examples: Inferring Visual Concepts from Image Sets

论文配图:Show Me Examples: Inferring Visual Concepts from Image Sets
图 1 · 摘自论文原文
  • 从多张图片中提取共享视觉概念,生成新图像
  • 在真实和合成数据上准确率提升,且能泛化到草图等新模态
  • 适合研究视觉推理、跨模态生成的学者

视觉语言模型(VLMs)可理解复杂文本指令,但在纯视觉上下文中推理能力不足。当前模型难以从一组示例图像中推断出共享概念,并将其应用于新输入。为此,我们提出视觉概念推断任务(VICIS):给定一个包含共享概念的小图像集和一张查询图像,模型需生成既保留该概念又与查询图像一致的新图像。实验表明,现有顶尖VLM在该任务上表现不佳,常忽略视觉上下文或产生有偏生成。为此,我们设计了一种训练框架与架构,能够从图像集中推断视觉概念,并从查询中提取特定概念嵌入。在合成数据及大规模ImageNet/WordNet数据上的实验显示,所提模型生成结果更准确、多样,且能泛化至未见概念和草图等新模态。

原文摘要 · Abstract (English)

Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, current models fail to infer shared concepts from sets of example images and apply them to new inputs. We introduce Visual Concept Inference from Sets (VICIS), a task that evaluates this capability. Given a small context set of images sharing a concept and a query image, the model must generate new images that preserve the context-defined concept while remaining consistent with the query. We show that state-of-the-art VLMs perform poorly on this task, often ignoring the visual context or defaulting to biased generations. To address this gap, we propose a training framework and architecture that learn to infer visual concepts from image sets and extract concept-specific embeddings from queries. Experiments on synthetic data and large-scale ImageNet/WordNet data show that our model generates more accurate and diverse outputs and generalizes to unseen concepts and modalities such as sketches.

视觉推理图像生成概念推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。