研究视觉语言模型跨模态记忆特性,发现知识迁移存在显著差距。
Quantifying Cross-Modality Memorization in Vision-Language Models
- 构建合成人物数据集,分模态训练并测试跨模态知识转移
- 单模态学习的知识在另一模态中回忆准确率有明显下降
- 该现象在强模型、去学习和多跳任务中均存在,适合多模态研究者参考
理解神经网络在训练过程中记住了什么及其方式至关重要,既关乎敏感信息的意外记忆,也影响真实世界知识密集型任务的有效知识获取。尽管以往研究主要关注单一模态的记忆,如大语言模型中的文本记忆或扩散模型中的图像记忆,但统一的多模态模型在实际应用中日益普及。本文聚焦于跨模态记忆的独特特性,系统研究视觉语言模型。为支持可控实验,我们首先引入一个包含多样化合成人物图像和文本描述的合成人物数据集。通过在单一模态上训练并在另一模态上评估,量化事实知识记忆与跨模态可转移性。结果表明,一个模态中学到的事实能转移到另一个模态,但在源模态与目标模态间回忆表现存在显著差距。此外,这种差距在多种场景下普遍存在,包括更强大的模型、机器去学习(machine unlearning)以及多跳推理情形。最后,我们提出一种基线方法以缓解此挑战。希望本研究能激发未来开发更鲁棒的多模态学习技术,提升跨模态迁移能力。
原文摘要 · Abstract (English)
Understanding what and how neural networks memorize during training is crucial, both from the perspective of unintentional memorization of potentially sensitive information and from the standpoint of effective knowledge acquisition for real-world, knowledge-intensive tasks. While previous studies primarily investigate memorization within a single modality, such as text memorization in large language models or image memorization in diffusion models, unified multimodal models are becoming increasingly prevalent in practical applications. In this work, we focus on the unique characteristics of cross-modality memorization and conduct a systematic study centered on vision-language models. To facilitate controlled experiments, we first introduce a synthetic persona dataset comprising diverse synthetic person images and textual descriptions. We quantify factual knowledge memorization and cross-modal transferability by training models on a single modality and evaluating their performance in the other. Our results reveal that facts learned in one modality transfer to the other, but a significant gap exists between recalling information in the source and target modalities. Furthermore, we observe that this gap exists across various scenarios, including more capable models, machine unlearning, and the multi-hop case. At the end, we propose a baseline method to mitigate this challenge. We hope our study can inspire future research on developing more robust multimodal learning techniques to enhance cross-modal transferability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。