arXiv:2501.04568cs.CVcs.AI2025-01被引 1

用少量人工图片+自生成文本,让视觉语言模型更准更少幻觉。

Feedback-Driven Vision-Language Alignment with Minimal Human Supervision

  • 仅需少量人工选图,通过自生成文本和反馈机制提升对齐
  • 图文生成准确率提升14%,物体召回率提高12%,幻觉显著减少
  • 适合资源有限但想优化模型性能的研究者

视觉语言模型(VLMs)在融合视觉与语言信息方面展现出巨大潜力,但其性能常受限于高质量图像-文本训练数据的大量需求。此类数据的筛选耗时且计算成本高。为此,我们提出SVP(基于采样的视觉投影)框架,无需依赖人工标注的图文对或偏好标注,即可增强视觉语言对齐。SVP利用少量人工选取的图像、自生成文本及预训练定位模型作为反馈机制,激发VLM中的潜在信息。我们在六个关键任务上评估该方法:图像描述、指代消解、视觉问答、多任务学习、幻觉控制和物体召回。结果表明,该方法实现显著提升,包括图像描述任务平均提升14%,物体召回率最高提升12%,幻觉大幅减少,同时保持问答能力。使用SVP后,小型VLM的幻觉水平接近五倍规模模型,初始指代能力差的模型性能提升超一倍,接近两倍大小模型的水平。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have demonstrated remarkable potential in integrating visual and linguistic information, but their performance is often constrained by the need for extensive, high-quality image-text training data. Curation of these image-text pairs is both time-consuming and computationally expensive. To address this challenge, we introduce SVP (Sampling-based Visual Projection), a novel framework that enhances vision-language alignment without relying on manually curated text-image pairs or preference annotation. SVP leverages a small set of manually selected images, self-captioning and a pre-trained grounding model as a feedback mechanism to elicit latent information in VLMs. We evaluate our approach across six key areas: captioning, referring, visual question answering, multitasking, hallucination control, and object recall. Results demonstrate significant improvements, including a 14 % average improvement in captioning tasks, up to 12 % increase in object recall, and significantly reduced hallucinations, while maintaining question-answering capabilities. Using SVP, a small VLM achieves hallucination reductions similar to a model five times larger, while a VLM with initially poor referring capabilities more than doubles its performance, approaching parity with a model twice its size.

视觉语言少样本幻觉控制自生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。