构建首个实时视觉知识评估数据集,解决大模型过时问题
Seeking and Updating with Live Visual Knowledge
- 设计LiveVQA数据集,覆盖12类实时视觉信息
- 模型在截止日期后内容上平均提升327%(用工具后)
- 探索参数高效微调,平衡适配器容量与模型能力
我们周围的世界持续变化,从实时新闻、社交媒体趋势到卫星影像和增强现实的基础设施更新。然而,多模态大语言模型(MLLMs)因训练数据存在固定截止日期而难以保持最新。为量化这一滞后现象,我们提出首个同类数据集LiveVQA,包含107,143个样本和12个类别,专门用于研究模型获取与更新实时视觉知识的能力。数据源自2024年4月至2025年5月的新闻文章、视频平台及学术出版物。对17个顶尖MLLMs的全面评测显示,模型在知识截止日期之后的内容上表现显著下降;使用工具或代理式视觉搜索框架可带来平均327%的性能提升。此外,我们探索了参数高效微调(PEFT)方法以更新模型的视觉知识,并深入分析了适配器容量与模型能力之间的关键平衡关系。所有实验数据与源代码已公开:https://livevqa.github.io。
原文摘要 · Abstract (English)
The visual world around us constantly evolves, from real-time news and social media trends to global infrastructure changes visible through satellite imagery and augmented reality enhancements. However, Multimodal Large Language Models (MLLMs), which automate many tasks, struggle to stay current, limited by the cutoff dates in their fixed training datasets. To quantify this stagnation, we introduce LiveVQA, the first-of-its-kind dataset featuring 107,143 samples and 12 categories data specifically designed to support research in both seeking and updating with live visual knowledge. Drawing from recent news articles, video platforms, and academic publications in April 2024-May 2025, LiveVQA enables evaluation of how models handle latest visual information beyond their knowledge boundaries and how current methods help to update them. Our comprehensive benchmarking of 17 state-of-the-art MLLMs reveals significant performance gaps on content beyond knowledge cutoff, and tool-use or agentic visual seeking framework drastically gain an average of 327% improvement. Furthermore, we explore parameter-efficient fine-tuning (PEFT) methods to update MLLMs with new visual knowledge. We dive deeply to the critical balance between adapter capacity and model capability when updating MLLMs with new visual knowledge. All the experimental dataset and source code are publicly available at: https://livevqa.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。