arXiv:2506.05302cs.CV2025-06NeurIPS被引 46

PAM模型可一键识别、解释、描述并分割图像视频中的任意区域。

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

论文配图:Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
图 1 · 摘自论文原文
  • 用LLM增强SAM2,实现区域级视觉理解与多模态输出
  • 150万图像+60万视频区域标注,支持细粒度语义分析
  • 速度快1.2-2.4倍,适合实时应用的轻量级系统

我们提出感知一切模型(Perceive Anything Model, PAM),一种概念简洁且高效的图像与视频区域级视觉理解框架。该方法在强大分割模型SAM 2基础上集成大语言模型(LLMs),实现对象分割与多样化区域特定语义输出的同步生成,包括类别、标签定义、功能解释和详细描述。关键组件语义感知器(Semantic Perceiver)将SAM 2丰富的视觉特征(蕴含通用视觉、定位与语义先验)高效转换为适合LLM理解的多模态标记。为支持多粒度理解,我们还开发了专用的数据精炼与增强管道,构建了一个高质量数据集,包含150万图像与60万视频的区域-语义标注,涵盖新颖的区域级流式视频描述数据。PAM设计轻量高效,在多种区域理解任务中表现强劲,运行速度比先前方法快1.2至2.4倍,且显存占用更低,为实际应用提供了可行方案。我们认为该有效方法将成为未来区域级视觉理解研究的重要基线。

原文摘要 · Abstract (English)

We present Perceive Anything Model (PAM), a conceptually straightforward and efficient framework for comprehensive region-level visual understanding in images and videos. Our approach extends the powerful segmentation model SAM 2 by integrating Large Language Models (LLMs), enabling simultaneous object segmentation with the generation of diverse, region-specific semantic outputs, including categories, label definition, functional explanations, and detailed captions. A key component, Semantic Perceiver, is introduced to efficiently transform SAM 2's rich visual features, which inherently carry general vision, localization, and semantic priors into multi-modal tokens for LLM comprehension. To support robust multi-granularity understanding, we also develop a dedicated data refinement and augmentation pipeline, yielding a high-quality dataset of 1.5M image and 0.6M video region-semantic annotations, including novel region-level streaming video caption data. PAM is designed for lightweightness and efficiency, while also demonstrates strong performance across a diverse range of region understanding tasks. It runs 1.2-2.4x faster and consumes less GPU memory than prior approaches, offering a practical solution for real-world applications. We believe that our effective approach will serve as a strong baseline for future research in region-level visual understanding.

视觉理解多模态图像分割视频描述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。