arXiv:2510.18876cs.CVcs.AI2025-10被引 16

让多模态大模型能精准理解任意区域的上下文关系,支持自由提问和复杂推理。

Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs

  • 通过区域对齐特征重放技术,结合全局上下文提升区域感知精度
  • 在GAR-Bench上超越InternVL3-78B,实现跨区域关系建模与组合推理
  • 零样本迁移能力强,视频理解表现优于同类模型

尽管多模态大语言模型在整体理解上表现优异,但在复杂场景中难以捕捉密集细节及物体间关系。现有区域级模型通常仅关注孤立区域,忽略关键全局上下文。为此,我们提出Grasp Any Region(GAR),实现全面的区域级视觉理解。通过高效的RoI对齐特征重放技术,GAR具备:(1) 利用必要全局上下文进行精确感知;(2) 建模多个提示间的交互关系;由此自然支持(3) 针对任意区域的自由形式问题回答,推动范式从被动描述转向主动对话。我们构建了GAR-Bench,不仅更准确评估单区域理解,更关键的是衡量多区域间交互与复杂推理能力。大量实验表明,GAR-1B在保持领先图文描述性能的同时(如在DLC-Bench上比DAM-3B高4.5分),在多提示关系建模与高级理解方面表现卓越,甚至在GAR-Bench-VQA上超过InternVL3-78B。更重要的是,零样本GAR-8B在VideoRefer-BenchQ上表现优于域内VideoRefer-7B,表明其强大能力可轻松迁移至视频任务。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) excel at holistic understanding, they struggle in capturing the dense world with complex scenes, requiring fine-grained analysis of intricate details and object inter-relationships. Region-level MLLMs have been a promising step. However, previous attempts are generally optimized to understand given regions in isolation, neglecting crucial global contexts. To address this, we introduce Grasp Any Region (GAR) for comprehen- sive region-level visual understanding. Empowered by an effective RoI-aligned feature replay technique, GAR supports (1) precise perception by leveraging necessary global contexts, and (2) modeling interactions between multiple prompts. Together, it then naturally achieves (3) advanced compositional reasoning to answer specific free-form questions about any region, shifting the paradigm from passive description to active dialogue. Moreover, we construct GAR-Bench, which not only provides a more accurate evaluation of single-region comprehension, but also, more importantly, measures interactions and complex reasoning across multiple regions. Extensive experiments have demonstrated that GAR-1B not only maintains the state-of-the-art captioning capabilities, e.g., outperforming DAM-3B +4.5 on DLC-Bench, but also excels at modeling relationships between multiple prompts with advanced comprehension capabilities, even surpassing InternVL3-78B on GAR-Bench-VQA. More importantly, our zero-shot GAR-8B even outperforms in-domain VideoRefer-7B on VideoRefer-BenchQ, indicating its strong capabilities can be easily transferred to videos.

多模态区域理解上下文建模推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。