对比视觉语言模型的记憶與檢索能力,發現微調模型更依賴記憶但準確率更高。
Quantifying Memorization and Parametric Response Rates in Retrieval-Augmented Vision-Language Models
- 透過錯誤檢索仍答對的案例,設計記憶度評估指標。
- 微調模型在WebQA上準確率72%,遠高於基礎模型的52%。
- 圖像問題的參數回應率比文字高15-25%,顯示模態差異。
大型語言模型在問答任務中表現出色,但衡量其對記憶與檢索的依賴程度的指標仍不成熟。儘管微調模型在封閉領域任務中表現最佳,通用模型如GPT-4o亦具備強大的零樣本能力。這引發了記憶、泛化與檢索之間的權衡問題。本文分析多模態檢索增強型視覺語言模型(VLMs)相比基線模型在訓練數據上的記憶程度。使用WebQA基準,比較微調模型與基線VLM在多跳檢索與問答任務上的表現,探討微調對數據記憶的影響。為量化端到端檢索與問答系統中的記憶行為,提出若干代理指標,通過分析檢索失敗但問答成功的案例。結果顯示,微調模型比檢索增強型VLM更依賴記憶,且準確率更高(WebQA測試集上72%對52%)。此外,首次實證比較文本與視覺模態的參數響應率,發現圖像問題的參數響應率始終比文字問題高出15-25%。研究結果提示未來工作需考慮不同模態間記憶差異,並協調記憶與泛化在聯合檢索-問答任務中的關係。
原文摘要 · Abstract (English)
Large Language Models (LLMs) demonstrate remarkable capabilities in question answering (QA), but metrics for assessing their reliance on memorization versus retrieval remain underdeveloped. Moreover, while finetuned models are state-of-the-art on closed-domain tasks, general-purpose models like GPT-4o exhibit strong zero-shot performance. This raises questions about the trade-offs between memorization, generalization, and retrieval. In this work, we analyze the extent to which multimodal retrieval-augmented VLMs memorize training data compared to baseline VLMs. Using the WebQA benchmark, we contrast finetuned models with baseline VLMs on multihop retrieval and question answering, examining the impact of finetuning on data memorization. To quantify memorization in end-to-end retrieval and QA systems, we propose several proxy metrics by investigating instances where QA succeeds despite retrieval failing. In line with existing work, we find that finetuned models rely more heavily on memorization than retrieval-augmented VLMs, and achieve higher accuracy as a result (72% vs 52% on WebQA test set). Finally, we present the first empirical comparison of the parametric effect between text and visual modalities. Here, we find that image-based questions have parametric response rates that are consistently 15-25% higher than for text-based questions in the WebQA dataset. As such, our measures pose a challenge for future work, both to account for differences in model memorization across different modalities and more generally to reconcile memorization and generalization in joint Retrieval-QA tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。