arXiv:2510.21182cs.CVcs.CL2025-10被引 2

动态评估多模态模型,防止数据污染与过拟合。

KBE-DME: Dynamic Multimodal Evaluation via Knowledge Enhanced Benchmark Evolution

  • 用图结构建模视觉问答,实现动态题库演化。
  • 可重构问题并引入外部知识,支持难度可控评估。
  • 适合评估多模态大模型真实能力,避免虚假性能提升。

多模态大语言模型(MLLMs)的快速发展亟需更可靠的评估方法。现有静态基准存在数据污染和饱和风险,导致性能评估虚高或误导。为此,我们首次将图结构用于表示静态或动态的视觉问答样本,提出知识增强型基准演化框架KBE。该框架先分析原始静态基准,再通过整合多模态知识,将其转化为可控制、可演化的动态版本。关键在于,KBE可基于原图重新选择视觉信息重构问题,也可结合外部文本知识扩展已有问题,实现探索程度可控的难度调节。大量实验表明,KBE有效缓解了数据污染与饱和问题,提供了对MLLM能力更全面的评估。

原文摘要 · Abstract (English)

The rapid progress of multimodal large language models (MLLMs) calls for more reliable evaluation protocols. Existing static benchmarks suffer from the potential risk of data contamination and saturation, leading to inflated or misleading performance evaluations. To address these issues, we first apply Graph formulation to represent a static or dynamic VQA sample. With the formulation, we propose Knowledge-enhanced Benchmark Evolution(KBE), a dynamic multimodal evaluation framework. KBE first analyzes the original static benchmark, then expands it by integrating multimodal knowledge, transforming the static benchmark into a controllable, dynamic evolving version. Crucially, KBE can both reconstruct questions by Re-selecting visual information in the original image and expand existing questions with external textual knowledge. It enables difficulty-controllable evaluation by adjusting the degree of question exploration. Extensive experiments demonstrate that KBE alleviates the risk of data contamination, data saturation, and provides a more comprehensive assessment of MLLM capabilities.

多模态评估动态基准知识增强MLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。