动态生成谜题数据集,让大模型评测永不落伍。
PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving
- 用自动生成机制构造全新谜题数据,避免数据污染。
- 构建含11840个样本的动态评测集,覆盖六大谜题任务。
- 适合关注模型长期评估与抗过拟合能力的研究者。
大型多模态模型(LMMs)在多种多模态任务中表现出色,但在现有静态评测基准上表现趋近饱和。这些基准常与预训练数据重叠,导致数据污染和复杂度固定。人工标注数据成本高且易受主观偏差影响。为此,我们提出完全动态的多模态评测框架Open-ended Visual Puzzle Generation(OVPG),通过原始素材采样、视觉内容生成和谜题规则设计三模块,自动生成新颖、多样、可验证的谜题数据。基于OVPG,我们构建了PuzzleBench,一个包含11,840个VQA样本的动态可扩展基准,涵盖六类精心设计的谜题任务,聚焦视觉识别、逻辑推理与上下文理解三大核心能力。该框架支持持续刷新数据集,适应LMMs演进,突破传统静态基准的局限性。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) have demonstrated impressive capabilities across a wide range of multimodal tasks, achieving ever-increasing performance on various evaluation benchmarks. However, existing benchmarks are typically static and often overlap with pre-training datasets, leading to fixed complexity constraints and substantial data contamination issues. Meanwhile, manually annotated datasets are labor-intensive, time-consuming, and subject to human bias and inconsistency, leading to reliability and reproducibility issues. To address these problems, we propose a fully dynamic multimodal evaluation framework, named Open-ended Visual Puzzle Generation (OVPG), which aims to generate fresh, diverse, and verifiable evaluation data automatically in puzzle-solving tasks. Specifically, the OVPG pipeline consists of a raw material sampling module, a visual content generation module, and a puzzle rule design module, which ensures that each evaluation instance is primitive, highly randomized, and uniquely solvable, enabling continual adaptation to the evolving capabilities of LMMs. Built upon OVPG, we construct PuzzleBench, a dynamic and scalable benchmark comprising 11,840 VQA samples. It features six carefully designed puzzle tasks targeting three core LMM competencies, visual recognition, logical reasoning, and context understanding. PuzzleBench differs from static benchmarks that quickly become outdated. It enables ongoing dataset refreshing through OVPG and a rich set of open-ended puzzle designs, allowing seamless adaptation to the evolving capabilities of LMMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。