PRIMA让视觉语言模型跨图像精准定位并推理,提升细粒度理解能力。
PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation
- 引入SQuARE模块,用查询式视觉令牌融合多图关系信息。
- 在M4SEG数据集上召回率和分割交并比分别提升7.83%和11.25%。
- 适合需要跨图对比分析的图像理解任务,如医疗影像、遥感监测。
尽管大型视觉语言模型(LVLM)能力显著提升,现有像素定位模型仍局限于单图场景,难以进行多图间的精细比较;而现有多图理解模型缺乏像素级定位。本文提出多图像像素定位推理任务及相应模型PRIMA,将像素级定位与多图推理结合,生成上下文丰富的像素级解释。核心是SQuARE视觉模块,在融合前通过紧凑的查询式视觉令牌注入跨图关系信息。为支持训练与评估,构建了新基准M4SEG,包含约74.4万条需跨图细粒度理解的问题-答案对。PRIMA在召回率和S-IoU上分别优于最先进基线7.83%和11.25%。消融实验验证了SQuARE模块在捕捉跨图关系上的有效性。
原文摘要 · Abstract (English)
Despite significant advancements in Large Vision-Language Models (LVLMs)' capabilities, existing pixel-grounding models operate in single-image settings, limiting their ability to perform detailed, fine-grained comparisons across multiple images. Conversely, current multi-image understanding models lack pixel-level grounding. Our work addresses this gap by introducing the task of multi-image pixel-grounded reasoning alongside PRIMA, an LVLM that integrates pixel-level grounding with robust multi-image reasoning to produce contextually rich, pixel-grounded explanations. Central to PRIMA is SQuARE, a vision module that injects cross-image relational context into compact query-based visual tokens before fusing them with the language backbone. To support training and evaluation, we curate M4SEG, a new multi-image reasoning segmentation benchmark consisting of $\sim$744K question-answer pairs that require fine-grained visual understanding across multiple images. PRIMA outperforms state-of-the-art baselines with $7.83\%$ and $11.25\%$ improvements in Recall and S-IoU, respectively. Ablation studies further demonstrate the effectiveness of the proposed SQuARE module in capturing cross-image relationships.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。