让视觉语言模型同时擅长图文检索和生成,且不互相干扰。
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
- 用对比学习与自回归建模联合训练,统一表示与生成任务。
- 在图文检索和生成任务上均达到当前最佳表现,有效减少物体幻觉。
- 适合需要兼具精准检索与高质量生成的多模态应用。
大型视觉语言模型(LVLM)在多模态任务中取得显著进展,能够跨视觉与文本域进行理解、推理与生成。尽管在生成任务中表现优异,现有LVLM在高保真表示学习任务(如生成用于检索的图像或文本嵌入)方面仍存在局限。近期研究尝试对LVLM进行表示学习微调,但该过程常导致生成能力下降。为解决这一权衡问题,本文提出CAFe——一种对比-自回归微调框架,通过将对比目标与自回归语言建模结合,使LVLM同时提升表示与生成能力。该方法在多模态检索与生成基准测试中均达到最先进水平,包括有效缓解物体幻觉(OH)。CAFe建立了一种新范式,实现单模型中嵌入与生成功能的协同优化,为未来兼具高精度检索与连贯生成能力的多模态模型奠定基础。
原文摘要 · Abstract (English)
The rapid advancement of large vision-language models (LVLMs) has driven significant progress in multimodal tasks, enabling models to interpret, reason, and generate outputs across both visual and textual domains. While excelling in generative tasks, existing LVLMs often face limitations in tasks requiring high-fidelity representation learning, such as generating image or text embeddings for retrieval. Recent work has proposed finetuning LVLMs for representational learning, but the fine-tuned model often loses its generative capabilities due to the representational learning training paradigm. To address this trade-off, we introduce CAFe, a contrastive-autoregressive fine-tuning framework that enhances LVLMs for both representation and generative tasks. By integrating a contrastive objective with autoregressive language modeling, our approach unifies these traditionally separate tasks, achieving state-of-the-art results in both multimodal retrieval and multimodal generative benchmarks, including object hallucination (OH) mitigation. CAFe establishes a novel framework that synergizes embedding and generative functionalities in a single model, setting a foundation for future multimodal models that excel in both retrieval precision and coherent output generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。