让模型根据文字或图文混合提示,精准定位多个分割掩码。
Refer to Any Segmentation Mask Group With Vision-Language Prompts
- 基于掩码中心的多模态大模型,实现复杂跨模态理解。
- 在200万条数据上训练,显著优于经典分割任务表现。
- 适合需要灵活交互的智能图像编辑与视觉问答场景。
近期图像分割模型已能生成高质量视觉实体掩码,但难以基于文本和视觉联合提示进行综合语义理解,限制了其在用户友好型交互应用中的效果。为此,我们提出全新的跨模态指代表达分割(ORES)任务:模型需根据纯文本或文本加参考视觉实体的任意提示,生成一组掩码。为解决该挑战,我们提出RAS框架,通过掩码中心的大规模多模态模型增强分割模型的复杂多模态交互与理解能力。为训练与评测ORES模型,我们构建了两个数据集:MaskGroups-2M(含200万条样本)和MaskGroups-HQ,涵盖多种由文本及参考实体定义的掩码组。大量实验表明,RAS在新提出的ORES任务以及经典指代表达分割(RES)和广义指代表达分割(GRES)任务中均取得卓越性能。
原文摘要 · Abstract (English)
Recent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehensive semantic understanding for complex queries based on both language and vision. This limitation reduces their effectiveness in applications that require user-friendly interactions driven by vision-language prompts. To bridge this gap, we introduce a novel task of omnimodal referring expression segmentation (ORES). In this task, a model produces a group of masks based on arbitrary prompts specified by text only or text plus reference visual entities. To address this new challenge, we propose a novel framework to "Refer to Any Segmentation Mask Group" (RAS), which augments segmentation models with complex multimodal interactions and comprehension via a mask-centric large multimodal model. For training and benchmarking ORES models, we create datasets MaskGroups-2M and MaskGroups-HQ to include diverse mask groups specified by text and reference entities. Through extensive evaluation, we demonstrate superior performance of RAS on our new ORES task, as well as classic referring expression segmentation (RES) and generalized referring expression segmentation (GRES) tasks. Project page: https://Ref2Any.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。