arXiv:2411.17606cs.CV2024-11被引 39

用大模型实现图像视频通用分割,能理解复杂指令

HyperSeg: Towards Universal Visual Segmentation with Large Language Model

  • 基于视觉大模型构建通用分割框架,支持图像视频统一处理
  • 在复杂推理任务上超越现有方法,实现精准细粒度语义理解
  • 适合需要强推理能力的多模态视觉应用开发者

本文旨在利用视觉大语言模型(VLLM)的强大推理能力,解决图像与视频感知中的通用分割问题。尽管当前统一分割方法取得进展,但在图像与视频场景适应性及复杂推理分割方面仍存在局限,难以应对多样化挑战性指令并准确理解细粒度视觉-语言关联。我们提出HyperSeg,首个基于VLLM的像素级图像与视频通用分割模型,涵盖通用分割任务和需强大推理能力与世界知识的复杂感知任务。为充分挖掘VLLM的识别能力与细粒度视觉信息,HyperSeg引入混合实体识别与细粒度视觉感知模块,结合时间适配器,实现对时序信息的全面理解。实验验证了该方法在通用图像与视频分割任务上的有效性,包括更复杂的推理感知任务。代码已开源。

原文摘要 · Abstract (English)

This paper aims to address universal segmentation for image and video perception with the strong reasoning ability empowered by Visual Large Language Models (VLLMs). Despite significant progress in current unified segmentation methods, limitations in adaptation to both image and video scenarios, as well as the complex reasoning segmentation, make it difficult for them to handle various challenging instructions and achieve an accurate understanding of fine-grained vision-language correlations. We propose HyperSeg, the first VLLM-based universal segmentation model for pixel-level image and video perception, encompassing generic segmentation tasks and more complex reasoning perception tasks requiring powerful reasoning abilities and world knowledge. Besides, to fully leverage the recognition capabilities of VLLMs and the fine-grained visual information, HyperSeg incorporates hybrid entity recognition and fine-grained visual perceiver modules for various segmentation tasks. Combined with the temporal adapter, HyperSeg achieves a comprehensive understanding of temporal information. Experimental results validate the effectiveness of our insights in resolving universal image and video segmentation tasks, including the more complex reasoning perception tasks. Our code is available.

通用分割视觉大模型视频理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。