提出首个图像语义复杂度评估任务与数据集,用语言模型辅助判断图像讲故事能力。
Is Your Image a Good Storyteller?
- 用语言模型分析图像语义丰富度,解决视觉任务中语义复杂度评估难题。
- 构建首个ISA数据集,验证方法在真实图像上有效提升语义评估精度。
- 适合关注认知科学、AI生成内容与跨文化视觉理解的研究者。
在实体层面量化图像复杂度较为直接,但语义复杂度的评估长期被忽视。事实上,不同图像间存在显著语义复杂度差异。语义丰富的图像能讲述生动引人的故事,应用场景广泛。例如,《Cookie Theft》这类图像因高语义复杂度,被广泛用于评估人类语言与认知能力。此外,语义丰富的图像对视觉模型发展也至关重要,因为语义贫乏的图像正逐渐失去挑战性。然而,此类图像稀缺,亟需更多类似《Cookie Theft》的图像以覆盖不同文化背景和时代。当前语义复杂度评估依赖人工专家与实证研究。因此,自动化评估图像语义丰富度成为挖掘或生成高质量图像的第一步,可推动认知评估、人工智能及其他应用发展。为此,我们提出图像语义评估(ISA)任务,构建首个ISA数据集,并提出一种利用语言模型解决视觉问题的新方法。在该数据集上的实验验证了方法的有效性。
原文摘要 · Abstract (English)
Quantifying image complexity at the entity level is straightforward, but the assessment of semantic complexity has been largely overlooked. In fact, there are differences in semantic complexity across images. Images with richer semantics can tell vivid and engaging stories and offer a wide range of application scenarios. For example, the Cookie Theft picture is such a kind of image and is widely used to assess human language and cognitive abilities due to its higher semantic complexity. Additionally, semantically rich images can benefit the development of vision models, as images with limited semantics are becoming less challenging for them. However, such images are scarce, highlighting the need for a greater number of them. For instance, there is a need for more images like Cookie Theft to cater to people from different cultural backgrounds and eras. Assessing semantic complexity requires human experts and empirical evidence. Automatic evaluation of how semantically rich an image will be the first step of mining or generating more images with rich semantics, and benefit human cognitive assessment, Artificial Intelligence, and various other applications. In response, we propose the Image Semantic Assessment (ISA) task to address this problem. We introduce the first ISA dataset and a novel method that leverages language to solve this vision problem. Experiments on our dataset demonstrate the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。