让图像区域描述更独特,支持多粒度精准刻画。
URECA: Unique Region Caption Anything
- 分阶段数据构建+MLLM生成,确保区域与描述一一对应
- 在自建数据集上达领先效果,且泛化至已有基准
- 适合需要细粒度图像理解的视觉语言任务
区域级描述旨在为图像特定区域生成自然语言描述,并突出其独特特征。然而,现有方法在多粒度下难以生成唯一性描述,限制了实际应用。为此,我们提出URECA数据集,一个面向多粒度区域描述的大规模数据集。不同于以往聚焦显著物体的数据集,URECA通过包含多样对象、部件和背景元素,确保区域与描述间唯一且一致的映射。核心是分阶段数据整理流程,每阶段逐步优化区域选择与描述生成。利用多模态大模型(MLLM)在各阶段生成更具区分性和上下文相关性的描述,提升准确率与语义多样性。基于该数据集,我们提出URECA模型,通过简单但有效的修改保留原有MLLM的空间属性(位置与形状),实现细粒度、语义丰富的区域描述。引入动态掩码建模与高分辨率掩码编码器,增强描述唯一性。实验表明,URECA在URECA数据集上达到当前最优性能,并在现有区域描述基准上具有良好泛化能力。
原文摘要 · Abstract (English)
Region-level captioning aims to generate natural language descriptions for specific image regions while highlighting their distinguishing features. However, existing methods struggle to produce unique captions across multi-granularity, limiting their real-world applicability. To address the need for detailed region-level understanding, we introduce URECA dataset, a large-scale dataset tailored for multi-granularity region captioning. Unlike prior datasets that focus primarily on salient objects, URECA dataset ensures a unique and consistent mapping between regions and captions by incorporating a diverse set of objects, parts, and background elements. Central to this is a stage-wise data curation pipeline, where each stage incrementally refines region selection and caption generation. By leveraging Multimodal Large Language Models (MLLMs) at each stage, our pipeline produces distinctive and contextually grounded captions with improved accuracy and semantic diversity. Building upon this dataset, we present URECA, a novel captioning model designed to effectively encode multi-granularity regions. URECA maintains essential spatial properties such as position and shape through simple yet impactful modifications to existing MLLMs, enabling fine-grained and semantically rich region descriptions. Our approach introduces dynamic mask modeling and a high-resolution mask encoder to enhance caption uniqueness. Experiments show that URECA achieves state-of-the-art performance on URECA dataset and generalizes well to existing region-level captioning benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。