arXiv:2410.20883cs.CV2024-10ECCV被引 13

不更新参数,让视觉语言模型自我集成,提升推理泛化能力。

Improving Generalization in Visual Reasoning via Self-Ensemble

  • 利用模型自身多视角输出进行自我集成,无需额外模型
  • 在SketchyVQA等任务上达到最新最优性能
  • 适合希望提升现有模型推理能力的开发者使用

视觉推理需要整合多模态感知与常识及外部世界知识。近年来大量大型视觉语言模型(LVLM)被提出,在多个领域和任务中展现出卓越的常识推理能力。然而,训练这些模型需消耗大量资源。近期方法不再从头训练,而是探索如何利用多种LVLM的能力,如集成方法。本文提出自集成(self-ensemble)方法,一种无需更新参数的训练自由方法,可提升模型的泛化与视觉推理能力。核心思想是:LVLM自身即可实现集成,无需依赖其他模型,从而释放其内部潜力。在多个基准测试中,该方法在SketchyVQA、外部知识VQA及分布外VQA任务上均取得当前最佳表现。

原文摘要 · Abstract (English)

The cognitive faculty of visual reasoning necessitates the integration of multimodal perceptual processing and commonsense and external knowledge of the world. In recent years, a plethora of large vision-language models (LVLMs) have been proposed, demonstrating outstanding power and exceptional proficiency in commonsense reasoning across diverse domains and tasks. Nevertheless, training such LVLMs requires a lot of costly resources. Recent approaches, instead of training LVLMs from scratch on various large datasets, focus on exploring ways to take advantage of the capabilities of many different LVLMs, such as ensemble methods. In this work, we propose self-ensemble, a novel method that improves the generalization and visual reasoning of the model without updating any parameters, a training-free method. Our key insight is that we realized that LVLM itself can ensemble without the need for any other LVLMs, which helps to unlock their internal capabilities. Extensive experiments on various benchmarks demonstrate the effectiveness of our method in achieving state-of-the-art (SOTA) performance on SketchyVQA, Outside Knowledge VQA, and out-of-distribution VQA tasks.

视觉推理自集成模型泛化LVLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。