arXiv:2411.00304cs.CVcs.MM2024-11NeurIPS被引 19

统一生成与判别训练,提升多模态大模型的语义理解与识别能力。

Unified Generative and Discriminative Training for Multi-modal Large Language Models

  • 通过动态时间对齐和新核函数,融合生成与判别训练机制。
  • 在生成与细粒度检索任务中均达到领先性能,尤其在认知类任务上表现突出。
  • 适合需要强语义区分与跨模态理解的研究者与应用开发者。

近年来,视觉语言模型(VLMs)主要采用两种训练范式。生成式训练使多模态大语言模型(MLLMs)能够处理复杂任务,但存在幻觉和弱物体判别问题。判别式训练(如CLIP)在零样本图像-文本分类与检索中表现优异,但在需要细粒度语义区分的复杂场景中表现不足。本文提出一种统一方法,整合两种范式的优点。将交错的图文序列作为输入通用格式,引入结构诱导训练策略,强化输入样本与MLLM隐状态间的语义关联,提升全局与细粒度语义捕捉能力。通过动态时间规整框架中的动态序列对齐,并引入新型核函数实现细粒度语义区分,有效平衡生成与判别任务。大量实验表明,该方法在多个生成任务中达到当前最优,尤其在需认知与判别能力的任务中表现显著。同时,在交错与细粒度检索任务上超越判别基准。通过检索增强生成策略,进一步提升部分生成任务性能,为视觉语言建模研究提供新方向。

原文摘要 · Abstract (English)

In recent times, Vision-Language Models (VLMs) have been trained under two predominant paradigms. Generative training has enabled Multimodal Large Language Models (MLLMs) to tackle various complex tasks, yet issues such as hallucinations and weak object discrimination persist. Discriminative training, exemplified by models like CLIP, excels in zero-shot image-text classification and retrieval, yet struggles with complex scenarios requiring fine-grained semantic differentiation. This paper addresses these challenges by proposing a unified approach that integrates the strengths of both paradigms. Considering interleaved image-text sequences as the general format of input samples, we introduce a structure-induced training strategy that imposes semantic relationships between input samples and the MLLM's hidden state. This approach enhances the MLLM's ability to capture global semantics and distinguish fine-grained semantics. By leveraging dynamic sequence alignment within the Dynamic Time Warping framework and integrating a novel kernel for fine-grained semantic differentiation, our method effectively balances generative and discriminative tasks. Extensive experiments demonstrate the effectiveness of our approach, achieving state-of-the-art results in multiple generative tasks, especially those requiring cognitive and discrimination abilities. Additionally, our method surpasses discriminative benchmarks in interleaved and fine-grained retrieval tasks. By employing a retrieval-augmented generation strategy, our approach further enhances performance in some generative tasks within one model, offering a promising direction for future research in vision-language modeling.

多模态大模型生成与判别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。