arXiv:2504.01954cs.CV2025-04被引 3

构建统一框架,让模型能精准定位从物体到部件的任意视觉目标。

Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities

  • 提出统一多粒度指代分割框架,融合对象与部件级任务。
  • 发布含3220万标注的MRES-32M数据集,支持细粒度视觉语言对齐。
  • UniRES++模型在多个基准上达领先性能,适合多粒度视觉理解研究者。

指代表达分割(RES)旨在分割与语言描述匹配的实体掩码。传统方法主要针对对象级定位,但现实场景需应对多层级目标粒度,如多对象、单对象或部件级参考。这带来描述多样性和语义细微差异的挑战。然而现有数据集与模型多聚焦于对象级定位,缺乏多粒度数据资源与统一框架。本文提出统一多粒度指代表达分割(MRES)任务,并发布包含部件级标注的RefCOCOm基准。同时构建了目前最大规模的视觉定位数据集MRES-32M,涵盖100万张图像中的3220万条掩码与描述。为应对多粒度挑战,提出UniRES++——一个整合对象级与部件级任务的统一多模态大模型,通过针对性设计实现细粒度视觉特征探索。该模型在多个基准上表现卓越,包括RefCOCOm(MRES)、gRefCOCO(通用RES)及RefCOCO/RefCOCO+/RefCOCOg(经典RES)。相关数据集与模型将公开于https://github.com/Rubics-Xuan/MRES,推动多粒度视觉语言研究发展。

原文摘要 · Abstract (English)

Referring expression segmentation (RES) aims at segmenting the entities' masks that match the descriptive language expression. While traditional RES methods primarily address object-level grounding, real-world scenarios demand a more versatile framework that can handle multiple levels of target granularity, such as multi-object, single object or part-level references. This introduces great challenges due to the diverse and nuanced ways users describe targets. However, existing datasets and models mainly focus on designing grounding specialists for object-level target localization, lacking the necessary data resources and unified frameworks for the more practical multi-grained RES. In this paper, we take a step further towards visual granularity unified RES task. To overcome the limitation of data scarcity, we introduce a new multi-granularity referring expression segmentation (MRES) task, alongside the RefCOCOm benchmark, which includes part-level annotations for advancing finer-grained visual understanding. In addition, we create MRES-32M, the largest visual grounding dataset, comprising over 32.2M masks and captions across 1M images, specifically designed for part-level vision-language grounding. To tackle the challenges of multi-granularity RES, we propose UniRES++, a unified multimodal large language model that integrates object-level and part-level RES tasks. UniRES++ incorporates targeted designs for fine-grained visual feature exploration. With the joint model architecture and parameters, UniRES++ achieves state-of-the-art performance across multiple benchmarks, including RefCOCOm for MRES, gRefCOCO for generalized RES, and RefCOCO, RefCOCO+, RefCOCOg for classic RES. To foster future research into multi-grained visual grounding, our RefCOCOm benchmark, MRES-32M dataset and model UniRES++ will be publicly available at https://github.com/Rubics-Xuan/MRES.

指代分割多粒度视觉语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。