arXiv:2504.09620cs.CLcs.AI2025-04被引 7

让多个视觉语言模型通过对话式学习互相提升生成能力

Metropolis-Hastings Captioning Game: Knowledge Fusion of Vision Language Models via Decentralized Bayesian Inference

  • 采用去中心化贝叶斯推理,让两个模型轮流描述图像并相互学习
  • 在无参考评价指标上实现稳定性能提升,跨数据集训练模型共享词汇
  • 适合研究多模型协作、知识融合与视觉语言理解的学者

我们提出元胞哈斯汀斯图像描述游戏(MHCG),通过一种类似语言游戏的机制,实现多个视觉语言模型(VLMs)之间的知识融合。现有融合方法存在推理开销大和架构限制问题,而MHCG通过去中心化贝叶斯推理规避了这些问题。该过程由两个预训练于不同数据集的VLM代理交替对图像进行描述,并从对方输出中学习。实验一表明,MHCG在无参考评估指标上取得一致提升;实验二分析生成描述中词汇的出现频率,揭示其促进模型间类别级词汇共享的能力。

原文摘要 · Abstract (English)

We propose the Metropolis-Hastings Captioning Game (MHCG), a method to fuse knowledge of multiple vision-language models (VLMs) by learning from each other. Although existing methods that combine multiple models suffer from inference costs and architectural constraints, MHCG avoids these problems by performing decentralized Bayesian inference through a process resembling a language game. The knowledge fusion process establishes communication between two VLM agents alternately captioning images and learning from each other. We conduct two image-captioning experiments with two VLMs, each pre-trained on a different dataset. The first experiment demonstrates that MHCG achieves consistent improvement in reference-free evaluation metrics. The second experiment investigates how MHCG contributes to sharing VLMs' category-level vocabulary by observing the occurrence of the vocabulary in the generated captions.

多模型融合视觉语言模型贝叶斯推理知识共享

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。