用视觉语言自举动态生成测评样本,解决模型评估过时与数据污染问题。
Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping
- 通过图像与语言协同修改生成新问答样本,保持语义一致性。
- 在SEEDBench、MMBench等基准上显著降低数据泄露风险。
- 适合评估持续进化的视觉语言模型,尤其关注泛化能力的研究者。
大型视觉语言模型(LVLM)在视觉感知与推理等多模态任务中表现出色,但在现有评估基准上仍面临静态性与预训练数据重叠的问题,导致评估复杂度固定且存在数据污染。为此,本文提出动态多模态评估协议——视觉语言自举(VLB),通过多模态自举模块动态生成新的视觉问答样本,同时利用判别模块确保新样本与原数据语义一致。通过组合不同自举策略,VLB可构建具有多样化复杂度的动态基准,使评估随模型能力演进而更新。在SEEDBench、MMBench和MME等多个基准上的实验表明,VLB有效降低了数据污染,并揭示了LVLM的真实性能瓶颈。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across multimodal tasks such as visual perception and reasoning, leading to good performance on various multimodal evaluation benchmarks. However, these benchmarks keep a static nature and overlap with the pre-training data, resulting in fixed complexity constraints and data contamination issues. This raises the concern regarding the validity of the evaluation. To address these two challenges, we introduce a dynamic multimodal evaluation protocol called Vision-Language Bootstrapping (VLB). VLB provides a robust and comprehensive assessment for LVLMs with reduced data contamination and flexible complexity. To this end, VLB dynamically generates new visual question-answering samples through a multimodal bootstrapping module that modifies both images and language, while ensuring that newly generated samples remain consistent with the original ones by a judge module. By composing various bootstrapping strategies, VLB offers dynamic variants of existing benchmarks with diverse complexities, enabling the evaluation to co-evolve with the ever-evolving capabilities of LVLMs. Extensive experimental results across multiple benchmarks, including SEEDBench, MMBench, and MME, show that VLB significantly reduces data contamination and exposes performance limitations of LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。