arXiv:2507.18915cs.CLcs.CV2025-07

从图像中挖掘视觉关联,生成越来越抽象的创意描述

Mining Contextualized Visual Associations from Images for Creativity Understanding

  • 通过分析图像中显著元素的上下文关联,自动生成创意描述
  • 构建了包含170万条创意描述的MSCOCO数据集,且抽象程度可调
  • 适合研究创意理解、视觉语言模型或艺术生成的学者使用

理解他人的创造性输出需要共享的联想语言。然而,训练如CLIP等视觉-语言模型时,通常依赖网络爬取的短句、以字面描述为主的替代文本数据集。本文提出一种方法,可从任意无标签图像数据集中挖掘显著视觉元素的上下文化关联,并支持任意规模扩展。给定图像,该方法能生成具有不同程度抽象性的高质量创意标题。我们基于此构建了新的视觉关联数据集和170万条来自MSCOCO图像的创意标题。人工评估表明,这些标题保持视觉一致性的同时,抽象程度明显提升。此外,在该数据集上微调视觉编码器,显著提升了在诗歌与隐喻可视化两个创造性领域中的零样本图文检索性能。我们已公开数据集、生成代码及模型供社区使用。

原文摘要 · Abstract (English)

Understanding another person's creative output requires a shared language of association. However, when training vision-language models such as CLIP, we rely on web-scraped datasets containing short, predominantly literal, alt-text. In this work, we introduce a method for mining contextualized associations for salient visual elements in an image that can scale to any unlabeled dataset. Given an image, we can use these mined associations to generate high quality creative captions at increasing degrees of abstraction. With our method, we produce a new dataset of visual associations and 1.7m creative captions for the images in MSCOCO. Human evaluation confirms that these captions remain visually grounded while exhibiting recognizably increasing abstraction. Moreover, fine-tuning a visual encoder on this dataset yields meaningful improvements in zero-shot image-text retrieval in two creative domains: poetry and metaphor visualization. We release our dataset, our generation code and our models for use by the broader community.

创意理解视觉关联生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。