arXiv:2503.19361cs.CV2025-03AAAI被引 2

用自然语言自动描述图像集合,结合大模型与知识图谱提升准确性。

ImageSet2Text: Describing Sets of Images through Text

  • 基于大模型和视觉问答链,逐步提取图像子集的关键概念。
  • 构建结构化概念图,实现对大规模图像集的可靠摘要。
  • 适合需要图像集合语义理解的科研与工业场景。

在大规模视觉数据时代,理解图像集合是一项挑战性且重要的任务。为此,我们提出ImageSet2Text,一种自动为图像集合生成自然语言描述的新方法。该方法基于大语言模型、视觉问答链、外部词汇图谱以及基于CLIP的验证机制,通过迭代方式从图像子集中提取关键概念,并组织成结构化的概念图。我们进行了大量实验,评估生成描述在准确性、完整性和用户满意度方面的表现。同时通过消融实验、可扩展性分析和失败案例研究考察了方法行为。结果表明,ImageSet2Text将数据驱动的AI与符号化表示相结合,能可靠地为各类应用中的大规模图像集合生成高质量摘要。

原文摘要 · Abstract (English)

In the era of large-scale visual data, understanding collections of images is a challenging yet important task. To this end, we introduce ImageSet2Text, a novel method to automatically generate natural language descriptions of image sets. Based on large language models, visual-question answering chains, an external lexical graph, and CLIP-based verification, ImageSet2Text iteratively extracts key concepts from image subsets and organizes them into a structured concept graph. We conduct extensive experiments evaluating the quality of the generated descriptions in terms of accuracy, completeness, and user satisfaction. We also examine the method's behavior through ablation studies, scalability assessments, and failure analyses. Results demonstrate that ImageSet2Text combines data-driven AI and symbolic representations to reliably summarize large image collections for a wide range of applications.

图像描述多图像理解大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。