arXiv:2509.18094cs.CVcs.AI2025-09NeurIPS被引 36

统一视觉提示与像素级推理,实现精准指代与分割。

UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning

论文配图:UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
图 1 · 摘自论文原文
  • 输入视觉提示后按需生成掩码,支持细粒度定位。
  • 在10个任务上验证,新设计的PixelQA任务展现强泛化能力。
  • 适合需要像素级理解的视觉问答与图像分析场景。

近年来,大型多模态模型(LMMs)在整体图像和视频语言理解方面取得了显著进展。然而,对细粒度像素级理解能力的关注仍不足,模型需实现视觉信号与语言语义之间的像素级对齐。尽管已有研究将LMMs应用于区域级描述和指代表达分割等任务,但这些模型仅能独立完成指代或分割,无法整合细粒度感知能力进行视觉推理。为此,我们提出UniPixel,一种可灵活理解视觉提示并生成掩码驱动响应的大规模多模态模型。其核心在于无缝融合像素级感知与通用视觉理解能力:在推理时根据视觉提示生成相关掩码,并基于这些中间指针执行后续推理,从而实现像素级推理。我们的方法已在10个涵盖图像与视频中像素级指代/分割及对象中心理解的任务上得到验证。此外,我们还设计了一个新的PixelQA任务,联合要求指代、分割与问答,以检验方法的灵活性。

原文摘要 · Abstract (English)

Recent advances in Large Multi-modal Models (LMMs) have demonstrated their remarkable success as general-purpose multi-modal assistants, with particular focuses on holistic image- and video-language understanding. Conversely, less attention has been given to scaling fine-grained pixel-level understanding capabilities, where the models are expected to realize pixel-level alignment between visual signals and language semantics. Some previous studies have applied LMMs to related tasks such as region-level captioning and referring expression segmentation. However, these models are limited to performing either referring or segmentation tasks independently and fail to integrate these fine-grained perception capabilities into visual reasoning. To bridge this gap, we propose UniPixel, a large multi-modal model capable of flexibly comprehending visual prompt inputs and generating mask-grounded responses. Our model distinguishes itself by seamlessly integrating pixel-level perception with general visual understanding capabilities. Specifically, UniPixel processes visual prompts and generates relevant masks on demand, and performs subsequent reasoning conditioning on these intermediate pointers during inference, thereby enabling fine-grained pixel-level reasoning. The effectiveness of our approach has been verified on 10 benchmarks across a diverse set of tasks, including pixel-level referring/segmentation and object-centric understanding in images/videos. A novel PixelQA task that jointly requires referring, segmentation, and question answering is also designed to verify the flexibility of our method.

像素级理解视觉推理多模态模型指代分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。