arXiv:2505.17726cs.CVcs.AI2025-05被引 7

用对象中心的注意力机制,让多模态大模型更精准理解与生成图像细节。

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM

  • 基于槽注意力设计对象级视觉标记,捕捉局部细节与高层语义。
  • 在多个视觉语言任务中显著超越基线模型,尤其擅长细粒度理解与生成。
  • 首次实现在真实图像上使用对象中心槽注意力,适合视觉生成与理解研究者。

近年来,多模态大语言模型(MLLMs)已成为实现通用人工智能的关键路径。特别是视觉-语言MLLMs已能从多模态输入生成文本和视觉输出。这要求高效且可处理的图像标记,以支持输入与输出的统一建模。然而,现有方法通常仅捕捉全局抽象概念或均匀分割的图像块,限制了模型对对象级细节的理解与生成能力。为此,我们提出一种基于槽注意力(Slot Attention)的对象中心视觉标记方法,专为MLLM设计。该方法结合Q-Former编码器、扩散解码器与残差向量量化,生成离散化的槽标记,既能编码局部视觉细节,又保留高层语义,并与文本数据对齐,无缝融入大模型的统一下一个词预测框架。实验表明,所提出的Slot-MLLM在涉及局部细节理解与生成的多种视觉语言任务中显著优于以往视觉标记方法。本工作首次验证了在真实自然图像上使用对象中心槽注意力在MLLM中的可行性。

原文摘要 · Abstract (English)

Recently, multimodal large language models (MLLMs) have emerged as a key approach in achieving artificial general intelligence. In particular, vision-language MLLMs have been developed to generate not only text but also visual outputs from multimodal inputs. This advancement requires efficient image tokens that LLMs can process effectively both in input and output. However, existing image tokenization methods for MLLMs typically capture only global abstract concepts or uniformly segmented image patches, restricting MLLMs' capability to effectively understand or generate detailed visual content, particularly at the object level. To address this limitation, we propose an object-centric visual tokenizer based on Slot Attention specifically for MLLMs. In particular, based on the Q-Former encoder, diffusion decoder, and residual vector quantization, our proposed discretized slot tokens can encode local visual details while maintaining high-level semantics, and also align with textual data to be integrated seamlessly within a unified next-token prediction framework of LLMs. The resulting Slot-MLLM demonstrates significant performance improvements over baselines with previous visual tokenizers across various vision-language tasks that entail local detailed comprehension and generation. Notably, this work is the first demonstration of the feasibility of object-centric slot attention performed with MLLMs and in-the-wild natural images.

多模态视觉标记槽注意力生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。