arXiv:2503.13847cs.CVcs.AI2025-03被引 1

用混合马尔可夫逻辑网分离视觉描述中预训练与微调的影响

Disentangling Fine-Tuning from Pre-Training in Visual Captioning with Hybrid Markov Logic

  • 基于混合马尔可夫逻辑网络建模图像与文本符号知识关联
  • 在MSCOCO上验证,使用LLM的BLIP2模型微调影响较小
  • 适合研究多模态模型知识演化和可解释性的研究人员

多模态系统具有复杂的处理流程,在大规模数据集上进行预训练后,再针对特定任务(如视觉描述)进行微调。然而,由于预训练的影响,难以区分模型在微调过程中学到的新知识与已有知识。本文通过混合马尔可夫逻辑网络(HMLN)对训练样本进行建模,将从字幕中提取的符号知识与从图像中提取的视觉特征相联系。对于生成的字幕,我们利用概率推理量化训练样本对生成结果的影响。我们在MSCOCO数据集上评估了两种推理方法,适用于不同类型的字幕生成模型。结果显示,对于采用大语言模型(LLM)的BLIP2模型,其微调过程对已习得知识的影响较小,可能因其已具备更通用的视觉描述能力。

原文摘要 · Abstract (English)

Multimodal systems have highly complex processing pipelines and are pretrained over large datasets before being fine-tuned for specific tasks such as visual captioning. However, it becomes hard to disentangle what the model learns during the fine-tuning process from what it already knows due to its pretraining. In this work, we learn a probabilistic model using Hybrid Markov Logic Networks (HMLNs) over the training examples by relating symbolic knowledge (extracted from the caption) with visual features (extracted from the image). For a generated caption, we quantify the influence of training examples based on the HMLN distribution using probabilistic inference. We evaluate two types of inference procedures on the MSCOCO dataset for different types of captioning models. Our results show that for BLIP2 (a model that uses a LLM), the fine-tuning may have smaller influence on the knowledge the model has acquired since it may have more general knowledge to perform visual captioning as compared to models that do not use a LLM

多模态视觉描述可解释性概率模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。