arXiv:2412.19806cs.CVcs.HC2024-12NeurIPS被引 100

Vitron统一处理图像视频的理解、生成、分割与编辑,实现像素级精准控制。

Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing

  • 融合图文编码器与视觉专家模块,支持多任务统一处理。
  • 在22个数据集上覆盖12项视觉任务,性能全面领先。
  • 适合需要跨任务协同的视觉生成与编辑研究者使用。

近期视觉大语言模型发展迅速,但仍面临细粒度实例理解不足、图像视频支持不统一、任务覆盖不全等挑战。本文提出VITRON,一个面向静态图像与动态视频的通用像素级视觉LLM,支持理解、生成、分割与编辑四大核心任务。基于大语言模型主干,VITRON在前端集成图像、视频及像素级区域编码器,在后端采用先进视觉专家模块,实现从低级到高级的视觉任务全覆盖。为实现高效指令传递,提出结合离散文本指令与连续信号嵌入的混合方法;设计多种像素级时空视觉-语言对齐学习,提升细粒度感知能力;引入跨任务协同模块,挖掘任务无关的细粒度特征,增强任务间协同效应。在12项视觉任务、22个数据集上的实验表明,VITRON展现出卓越的综合性能,揭示了构建更统一多模态通用模型的巨大潜力。

原文摘要 · Abstract (English)

Recent developments of vision large language models (LLMs) have seen remarkable progress, yet still encounter challenges towards multimodal generalists, such as coarse-grained instance-level understanding, lack of unified support for both images and videos, and insufficient coverage across various vision tasks. In this paper, we present VITRON, a universal pixel-level vision LLM designed for comprehensive understanding, generating, segmenting, and editing of both static images and dynamic videos. Building on top of an LLM backbone, VITRON incorporates encoders for images, videos, and pixel-level regional visuals within its frontend modules, while employing state-of-the-art visual specialists as its backend, via which VITRON supports a spectrum of vision end tasks, spanning visual comprehension to visual generation, from low level to high level. To ensure an effective and precise message passing from LLM to backend modules for function invocation, we propose a novel hybrid method by simultaneously integrating discrete textual instructions and continuous signal embeddings. Further, we design various pixel-level spatiotemporal vision-language alignment learning for VITRON to reach the best fine-grained visual capability. Finally, a cross-task synergy module is advised to learn to maximize the task-invariant fine-grained visual features, enhancing the synergy between different visual tasks. Demonstrated over 12 visual tasks and evaluated across 22 datasets, VITRON showcases its extensive capabilities in the four main vision task clusters. Overall, this work illuminates the great potential of developing a more unified multimodal generalist. Project homepage: https://vitron-llm.github.io/

视觉LLM多任务统一像素级控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。