arXiv:2411.15411cs.CV2024-11CVPR被引 30

让AI精准描述图像任意区域的细粒度组合特征

FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity

  • 用任意分割掩码作为输入,支持多粒度区域描述
  • 在新数据集上优于现有模型,实现更准的组合语义匹配
  • 适合需要精细视觉理解的场景,如医疗影像分析

大视觉语言模型(VLMs)虽在多模态任务中表现优异,但在细粒度区域组合信息感知上存在不足,难以准确对齐分割掩码与语义,也难精确描述所指区域的组合属性。为解决此问题,我们提出FINECAPTION,一种可接受任意掩码作为参照输入、处理高分辨率图像的新型VLM,支持不同粒度下的组合图像描述。为此,我们构建了COMPOSITIONCAP数据集,专注于多粒度区域组合图像描述任务,引入属性感知的区域描述能力。实验表明,该模型在多个指标上优于当前先进VLMs。同时,我们分析了现有VLM在组合区域描述任务中的表现,揭示了其在视觉提示识别方面的短板,为未来VLM设计与训练提供了改进方向。

原文摘要 · Abstract (English)

The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate reasoning across various applications, including image and video captioning, visual question answering, and cross-modal retrieval. Despite their superior capabilities, VLMs struggle with fine-grained image regional composition information perception. Specifically, they have difficulty accurately aligning the segmentation masks with the corresponding semantics and precisely describing the compositional aspects of the referred regions. However, compositionality - the ability to understand and generate novel combinations of known visual and textual components - is critical for facilitating coherent reasoning and understanding across modalities by VLMs. To address this issue, we propose FINECAPTION, a novel VLM that can recognize arbitrary masks as referential inputs and process high-resolution images for compositional image captioning at different granularity levels. To support this endeavor, we introduce COMPOSITIONCAP, a new dataset for multi-grained region compositional image captioning, which introduces the task of compositional attribute-aware regional image captioning. Empirical results demonstrate the effectiveness of our proposed model compared to other state-of-the-art VLMs. Additionally, we analyze the capabilities of current VLMs in recognizing various visual prompts for compositional region image captioning, highlighting areas for improvement in VLM design and training.

图像描述细粒度理解视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。