arXiv:2506.00721cs.CVcs.LG2025-06被引 7

构建9.7万张含上下文错位图像的数据集,推动视觉场景理解研究。

Common Inpainted Objects In-N-Out of Context

  • 用扩散模型替换COCO图像中的物体,生成上下文一致与不一致的图像。
  • 通过大视觉语言模型验证,97,722张图均被准确分类为在/不在上下文中。
  • 支持细粒度上下文推理、从场景预测新物体、增强假图检测等任务。

我们提出COinCO数据集,解决现有视觉数据集中缺乏上下文错位样本的问题。通过基于扩散模型的方法系统替换COCO图像中的物体,生成97,722张独特的图像,涵盖上下文一致与不一致的场景,有效支持上下文学习。每个修复后的物体均经大视觉语言模型评估,精确分类为在或不在上下文中。该数据集支持三项关键任务:(1) 基于三个标准的细粒度上下文推理方法,判断物体是否合理;(2) 新的“从场景预测物体”任务,在实例与群组语义层面预测场景中应存在的新物体;(3) 在无需微调的情况下,提升现有方法的假图检测能力。COinCO提供可控的上下文变化测试平台,为计算机视觉中的上下文感知理解(包括图像取证)奠定基础。代码与数据集见https://co-in-co.github.io/。

原文摘要 · Abstract (English)

We present Common Inpainted Objects In-N-Out of Context (COinCO), a novel dataset addressing the scarcity of out-of-context examples in existing vision datasets. By systematically replacing objects in COCO images through diffusion-based inpainting, we create 97,722 unique images featuring both contextually coherent and inconsistent scenes, enabling effective context learning. Each inpainted object is meticulously verified and categorized as in- or out-of-context through Large Vision Language Model assessments. We demonstrate three key tasks enabled by COinCO: (1) a fine-grained context reasoning approach that classifies objects as in- or out-of-context based on three criteria; (2) a novel Objects-from-Context prediction task that determines which new objects naturally belong in given scenes at both instance and clique level semantics, and (3) context-enhanced fake detection on state-of-the-art methods without fine-tuning. COinCO provides a controlled testbed with contextual variations, establishing a foundation for advancing context-aware visual understanding in computer vision, including image forensics. Code and dataset are available at https://co-in-co.github.io/.

数据集上下文理解图像伪造检测扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。